The alert had gone quiet, which was the problem. A rollover that hits a duplicate-key error is supposed to fail loud and then heal itself on retry. Ours failed loud once, and then every retry failed the same way, forever, and nobody noticed because a customer only calls in once.

The bug

Season rollover runs as a sequence: create the new season’s records, then a cleanup pass that catches duplicate-key collisions left over from a prior partial run and clears them out. That ordering was backward. The cleanup step that clears duplicates ran after the step that creates them, not before. So the first time a rollover hit a duplicate key, it threw, and the retry hit the exact same duplicate, because nothing had run yet to remove it. The cleanup code existed. It was correct. It was just downstream of the failure it was supposed to prevent.

Once you see it, the failure mode is obvious: any site that hit this once could never rerun a rollover successfully again. Not “would need a manual fix eventually.” Permanently stuck, on every future attempt, until someone reordered the two steps.

Finding it

I went in expecting a data problem, something in the source rows for a specific site. That was the wrong turn, and it cost real time: I pulled the failing site’s season data, diffed it against a season that had rolled over cleanly, and found nothing structurally different. No orphaned rows, no schema drift, nothing that explained why one key collided and the other didn’t. I spent a while on that path because a duplicate-key error looks like a data problem from the outside. It’s a database engine telling you a value already exists, so the natural move is to go look at the values.

The actual answer was in the rollover script’s step order, not the data. The cleanup routine was defined, tested, and worked exactly as written; it just came after the create step in the sequence instead of before it. A duplicate key thrown at create time meant cleanup never got a chance to run. Every retry hit create first, before cleanup, so every retry recreated the same collision it needed cleanup to have already cleared.

The fix

Moved the cleanup pass ahead of the create step, so any leftover duplicate from a prior run gets cleared before the new season tries to write over it. Shipped to develop for 3.361.

Three customer sites were confirmed stuck by this. Only one had actually called in, which is the part that makes this worth writing down: two sites had been silently unable to complete their rollover and nobody knew, including us. I ran the fix against production for the two that were straightforward and completed their rollovers end to end.

The part I didn’t want to repeat was finding out about the other two by phone. So the last piece wasn’t the fix, it was a scan tool that walks sites looking for rollovers stuck in this specific failure state, so we find them before a customer does. That’s the actual lesson from the wrong turn earlier: a duplicate-key error reads like a symptom you can diagnose from the data, but the root cause was purely sequential, and no amount of staring at rows was going to surface a step-ordering bug. The scan doesn’t try to be clever about data either. It just checks whether a site is sitting in the specific stuck state this bug produces, because that’s the only thing worth checking for.