The clone-cleanup job was supposed to be routine: 30 repos, fast-forward the stale base branches, delete dead leftover branches, close out the tickets sitting open against branches that no longer matched them. Two hours and seventeen minutes, filed under maintenance. It turned up two bugs that were about to ship broken.
Why 30 clones exist at all
The platform runs as one codebase cloned per deployment, 30 of them, each tracking its own base branch and its own set of open tickets. That model works until the clones drift. A branch gets created against the vmta-routing ticket, the ticket gets reassigned or closed elsewhere, and the branch just sits there, still open, still technically mergeable, no longer pointing at anything real. Multiply that by 30 repos and you get a pile of branches that look active in git branch -a but aren’t tied to anything a human is tracking.
The audit’s actual job was mechanical: walk all 30 clones, diff each base branch against its ticket’s current status, fast-forward the ones that were just behind, delete the ones that were dead, and flag the ones where the branch and the ticket disagreed about what was supposed to happen.
The mismatch that wasn’t cosmetic
Most of the flagged mismatches were exactly what you’d expect: a branch for a ticket that had already merged somewhere else, safe to delete. But one mismatch was the vmta-routing ticket, sitting on a branch that looked stale by the ticket-tracking logic but was actually the only place the fix existed. The ticket metadata said “stale,” the code said “unmerged and needed.” Routing was picking the wrong vmta for a subset of outbound mail, silently, because nothing in the pipeline had exercised that path since whoever branched it. If the cleanup script had trusted the ticket status and deleted the branch, the fix would have vanished and nobody would have noticed until mail started bouncing from the wrong source again.
The second one, the push-guard ticket, was a fix that had stalled mid-flight. Same shape of failure: a branch that the naive drift check would have called dead, but that was actually carrying a change nothing else had picked up. Both got merged and shipped live in the same pass instead of getting deleted as noise.
What the audit script actually got wrong
The first pass at the cleanup script did what I described above: treat “branch behind base and ticket closed or reassigned” as the signal for delete. That’s the wrong signal. A branch can be stale by commit distance and still be the only copy of a real fix, if the ticket system and the branch went out of sync for reasons that have nothing to do with whether the code is needed. I ran the first sweep on that logic, and it flagged both the vmta-routing ticket and the push-guard ticket for deletion. I caught it before the delete step ran, only because I diffed the flagged branches against develop by hand before letting the script execute the delete, and the diffs weren’t empty like a truly dead branch’s diff would be. If I’d trusted the flag list, both fixes are gone and I don’t find out until the vmta routing bug resurfaces in a support ticket weeks later with no branch left to bisect from.
The fix to the script was to add a second check before any delete: does this branch’s diff against its base contain anything not already present in develop. Stale-by-ticket-status and stale-by-content-diff turned out to be two different conditions, and only the second one is safe to act on.
What came out of it
Two live fixes (the vmta-routing ticket and the push-guard ticket), a follow-up ticket to revive the KB redirect-follow UI, and a tooling-gap ticket against the clone-cleanup scripts themselves, since the diff-check wasn’t there when the audit started and needs to be permanent, not something I remember to do by hand next time.
The uncomfortable part isn’t that the scripts had a gap. It’s that “audit and clean up 30 clones” reads like busywork right up until the busywork surfaces a routing bug already live in production mail. The maintenance pass and the bug hunt weren’t two different tasks. They were the same task, and the only reason the bugs surfaced is that the cleanup forced someone to actually look at every branch instead of trusting what the ticket tracker said about it.