My co-founder flagged something small that turned out not to be small: a routine step could fail partway through, and when it did, the safety net built to catch broken accounts couldn’t see the damage anymore. The bug didn’t just break the account. It erased the one signal that would have told anyone it was broken.

The step in question did three separate things in a required order, restore one setting, restore another, then a final action that finished the job. If that last action failed, the code simply reported the failure and stopped. The first two changes had already gone through by then, so the account was left half fixed: some settings updated as if everything had worked, but the piece that actually mattered still missing. Worse, one of the settings that got changed early was the exact thing a separate, regular check used to find broken accounts in the first place. So a partial failure didn’t just leave a mess. It also removed the account from the list of things anyone was watching.

The obvious instinct is to reorder things so the risky step happens first, before anything else changes. That didn’t work here, because the steps genuinely depended on each other in that order. One setting had to be in place before the final action would even target the right record, and the same setting was also what made the record findable to the routine check afterward. Swapping the order would have traded a recoverable failure for a different, worse kind of mistake.

What shipped instead wasn’t a fancier retry or a bigger safety net. It was a plain undo: before making any change, record exactly what the account’s settings were beforehand. If the final action fails, put those settings back exactly as they were, not to some fresh default that would look like nothing had happened recently. That distinction mattered. A fresh-looking value would have quietly reset the clock on how long the account had been in trouble, making an old problem look brand new. Restoring the real original value kept the account visible to the exact same routine check that already knew how to find and fix it. No new detection process was built. The one that already existed just needed the truth handed back to it.

One more small thing went in alongside it: if the undo itself failed, that got said plainly instead of swallowed quietly. A partial fix that fails silently a second time is the same problem all over again, just one layer deeper.

The lesson generalizes past any one system. Any multi-step process that can fail partway, a form submission, a multi-part refund, an onboarding checklist, has the same choice hiding in it: when something breaks halfway through, does the failure leave a state that your existing checks can still recognize as broken, or does it accidentally erase the very signal that was supposed to bring a person back to fix it?

The part worth copying isn’t “add a rollback.” It’s the specific choice of what the rollback restored: the account’s own original setting, not a fresh-looking replacement, because that original value was the exact thing the existing check was already reading to decide what counted as broken. If a process in your business can fail halfway through, check that same detail before building anything new: when it fails, does it put back the specific value your monitoring already keys off, or does it quietly leave something that looks fine instead?