The retro scanner reported 46 recurring failures. Its next run reported 33. Nothing in the underlying work had improved between those two numbers. The counter had been wrong.
I use retros to turn repeated agent and workflow failures into corrections. The loop only helps if the measurement answers a narrow question: what did we actually have to fix because the same problem kept returning? On this pass, the scanner was answering a different question without saying so. It was counting routine skill boilerplate changes as corrections.
That made the report worse than useless. It gave me a precise number that looked like evidence, while mixing maintenance work with the failures I was trying to understand.
The number that looked plausible
The bug did not announce itself as a crash or an empty result. The scanner worked. It found changes. It produced a total. A total of 46 even had the uncomfortable shape of a real result, large enough to suggest that our process was leaking the same mistakes over and over.
But a change to standard skill boilerplate is not necessarily a correction. It can be routine maintenance of the instructions that support the system. Treating it as a response to a recurring failure makes the metric count activity, not learning.
That distinction is easy to blur when a retro system starts from changed material. A change can tell you that something was edited. It cannot, by itself, tell you why. The scanner had collapsed two distinct states into one bucket. A change made after a recurring failure is evidence of a pattern worth tracking. Routine skill boilerplate is maintenance. It is not evidence of that pattern.
The count was therefore not a count of 46 failures. It was a count of 46 things that satisfied a bad definition.
I checked the gauge before fixing the process
The retro-driven pass still found six real recurring failure patterns to fix. That work was valuable. It also exposed the problem with the scanner because the reported total did not cleanly describe the corrections I could account for.
The diagnostic path was not complicated. I held the actual work constant and inspected the category the scanner was assigning to each hit. The routine boilerplate was inflating the same counter as the real corrections. The failure was in classification, not in the retro itself.
I considered the easy escape first: leave the rule alone and treat 46 as a rough directional signal. That loses the whole point of the scanner. A noisy total cannot tell me whether a failure pattern is returning, whether the correction rate is falling, or whether a batch of maintenance just made the graph jump.
I also could have lowered the threshold or hidden the boilerplate after the fact. Neither would repair the definition. It would only make the number look more comfortable. The fix had to separate an actual correction from routine skill boilerplate at the point where the scanner decides what belongs in the report.
After that change, the same retrospective read 33. The difference of 13 was not a victory lap. It was the size of the measurement error.
The revised result is less flattering and more useful
A drop from 46 to 33 normally sounds like improvement. Here, reading it that way would recreate the bug in prose. We had not suddenly eliminated 13 recurring failures. We had stopped attributing 13 routine boilerplate changes to failures in the first place.
That gives the remaining 33 more weight, not less. They are closer to the set of changes that should drive follow-up work, revisions to the operating skills, and another retro later. I logged two follow-up items from the pass, including the scanner defect itself, because a reporting bug inside the feedback loop is production work. It changes what we believe about the system.
The useful rule is not that every retro needs a perfect score. It is that a score needs a defensible noun. Before I use a number to tell me what keeps going wrong, I now ask what each row in that number actually means.
A scanner that counts boilerplate as learning does not just add noise. It teaches the system a false story about where the work is.