A production box needed a reboot it had been putting off for weeks. Nothing dramatic, just a backlog of staged operating-system patches that don’t take effect until the machine restarts. The plan for that reboot went through two rounds of correction before it ran, and a third round immediately after, because the box didn’t behave the way any version of the plan expected.

The first number was a guess wearing a fact’s clothes

The earliest draft of the maintenance plan stated the obvious-sounding detail: the site would be down for about three minutes during the reboot. It read like a fact because that’s how routine downtime windows get written. It was closer to a placeholder, a number nobody had actually checked against this specific box.

Re-checking the box directly, ahead of actually scheduling the window, turned up the real picture: a two-month backlog of staged patches, several rounds deep, still waiting to apply. A box with that much staged work can sit mid-reboot applying updates for a long stretch on both the way down and the way back up. The three-minute figure got struck and replaced with a wide range, twenty to forty-five minutes, along with a blunt note about what happens if the old number survives into the actual window: whoever’s watching stands there at the twenty-minute mark, no clear sign anything’s wrong, wondering whether the box died.

The window ran, and the plan was still incomplete

The reboot happened before dawn. An observer watched from outside the box the whole time: a snapshot beforehand, the site polled every minute, everything logged. The actual downtime came in under the corrected range, roughly two minutes. Then, about twenty minutes after the box came back and everything looked finished, it went down again, on its own, for another two minutes.

Nothing in either version of the plan mentioned a second reboot. It turned out to be entirely expected: a separate system component finishes installing the staged patches after the first restart completes, and queues its own follow-up restart to finish the job. That’s normal, documented behavior once you know to look for it. Nobody had looked for it, because nobody had run this specific box through this specific procedure before. The plan had a shape for “the reboot,” singular, and reality had two.

What the correction actually fixed

The maintenance document got rewritten the same day, in four places. It now says explicitly to expect a second, unannounced restart, and how to tell it apart from a real failure by reading the log entry that names what triggered it. The timing range widened again, with the one real measurement attached as a data point rather than a replacement guess. A resource ceiling that had quietly shifted on its own across the reboot, with nobody touching its configuration, got documented so the next comparison doesn’t read it as damage. And a handful of warning-level log entries that fire on every single boot of that box, before and after, got listed as known noise, so the next person watching doesn’t spend time chasing something that was never new.

Why the first number stuck around as long as it did

Nobody who wrote “about three minutes” believed it was solid. It was a first pass, meant to be checked. The problem is that a number sitting in an operational document doesn’t carry a visible confidence level. It reads exactly the same whether it came from a real measurement or from someone’s best guess on the first draft, and the next reader, or the same person a few days later under time pressure, has no way to tell the difference just by looking at it. It took a direct re-check against the actual box, not another read of the doc, to catch that the number was wrong. And it took an actual run of the real procedure, not another round of reasoning about it, to find the second reboot that even the corrected plan had missed.

A written estimate is a claim about the world, and claims stop being trustworthy the moment they’re carried forward without being re-checked against the thing they’re claiming something about.