The alert wasn’t an alert. It was five rows that had quietly moved off the product board and onto the internal one, and I only caught it because I happened to open the internal board first that morning.
The classifier is a nightly job: pull everything touched in the last cycle, look at where the commits landed, and file each ticket into either the product backlog or the internal ops backlog. Five tickets that week were customer-facing work, TLS cert renewal, a visibility bug on a customer’s tournament page, downgrade billing logic, all filed as internal. I reversed three of them within minutes of looking. The other two took longer because the shape of the mistake wasn’t obvious yet.
What the rule actually said
The classifier’s logic was repo-based: if the commit shipped from our internal tooling repo, tag it internal; if it shipped from the product API repo or the main product repo, tag it product. That’s a reasonable first pass, and it had been running clean for weeks because most internal-repo commits really were internal work, agent hooks, guardrails, dashboards.
Then the ticket tracking the failing tournament images came through. The actual work was: customer sites were failing to load images, root cause was an expired TLS cert on the media origin server, fix was reinstalling a replacement cert that had been issued back in May and never deployed. That’s customer-facing breakage with a same-day fix. But the commit that closed it out lived in our internal tooling repo, because that’s where our infra scripts live, so the classifier filed it as internal tooling maintenance. Same story for the downgrade/auto-renew billing fix, that one had a customer being charged a full year early, and it got filed next to a lint-error cleanup ticket because the patch touched an internal billing helper.
The wrong turn
My first fix was to patch the individual rows. I went into the ticket tracker, manually re-filed the three tickets I’d caught, and told myself I’d spot-check the nightly output more carefully going forward. That lasted about four days. The next batch, two more tickets slipped through the same way, one of them the tournament bracket visibility bug, which is unambiguously a customer-facing defect but shipped as a one-line config read inside our internal tooling repo because that’s where the diagnostic tooling lived.
Re-filing rows doesn’t fix anything. It’s a correction to the output, not to the machine that produced the output, so the same misclassification just regenerates itself on the same schedule. I’d spent two review cycles doing manual triage that the classifier should have done itself, and the backlog was no more trustworthy at the end of the second cycle than the first.
The actual fix
The question the rule needed to answer wasn’t “which repo did the commit land in,” it was “who does this deliver to.” So I rewrote the classifier to look at the ticket’s own delivery target instead of the commit’s file path: does the fix change behavior a customer or their site visitors experience, regardless of which repo housed the code that changed it. Concretely, that meant checking the ticket description and linked customer/site references first, and falling back to repo path only when no delivery target was named at all, which in practice is true for genuinely internal work like the Claude settings hook migration or the CI lint fix.
I ran it back against the prior week’s 71 sessions and 27 commits to check for regressions. The infra-only tickets (hook migration, CI debugging, prod reboot triage) stayed internal. The five customer-facing ones, cert renewal, billing fix, tournament visibility, all classified correctly on the first pass, no manual re-filing.
Why the row-fix wasn’t good enough twice
This was the second time that week a review turned into a rewrite of the rule instead of a fix to its output. The classifier had also been running unattended for hours at a stretch (the log shows the same untitled session firing every half hour overnight). It’s easy to let that cadence train you into treating a misfile as a one-off glitch to nudge back into place. But an unattended job that’s wrong in a structural way will produce the same wrong answer at the same time tomorrow. The value of catching it by hand wasn’t the three rows I moved back, it was noticing that the classification question itself was the wrong question, and that only shows up if you’re actually reading what the tickets say, not just where the diff landed.