I had just finished cleaning up a real problem on a live system, and I made a mistake trying to prove my own safety net was safe.

The work itself was routine: a fix to stop a background process from creating orphaned records, plus a cleanup of the ones that had already piled up. Standard practice for a change like that is to write a paired plan for undoing it, in case anything needs to be reversed later. I wrote one. Then I tried to double-check that plan before ever running it for real, and that’s where things actually went wrong.

A safety check that wasn’t checking anything

I wanted to confirm my undo plan was written correctly without actually running it against anything, so I turned on a feature built for exactly that: look over the steps, don’t execute them. I ran it against the live system, treating that as completely harmless.

It wasn’t. That feature suppresses execution for one particular kind of step. Everything else runs exactly as if the check had never been turned on, and most of an undo plan is made of exactly that “everything else.” So the “safety check” ran the plan. For real. Something I’d just finished setting up got undone, and the cleanup I’d just completed reverted right back to its messy state, as if none of it had happened.

Catching it in the same breath

The only reason this didn’t turn into a real incident is that I was already running verification checks as part of the same pass, and they came back wrong immediately. There was no gap between the mistake happening and someone noticing. Everything that had been undone got restored properly, the cleanup got redone, and the end state got checked against what it was supposed to look like. It matched.

While putting it back together, I found something else: the undo plan itself had a second problem sitting in it that hadn’t fired yet, because nothing had needed the real undo before. A step in it assumed nothing new had been added since the original change ran, and that assumption was already wrong the same day. I fixed that too, before it ever had a chance to matter for real.

The actual mistake

I didn’t skip a safety step. I added one specifically to avoid running something risky before I was sure it was right. What I got wrong was trusting what the feature’s name implied instead of checking what it actually covers. “Won’t execute anything” sounds like a blanket promise. It’s a promise about one specific case and silent about everything else, including the exact kind of step my plan was mostly made of.

The lesson isn’t “don’t use safety modes.” It’s that a feature’s name is a claim, and a claim about safety is exactly the kind of thing worth testing somewhere that isn’t live before trusting it somewhere that is. I had the discipline to write an undo plan and try to check it before use. I didn’t have the discipline to confirm my “won’t execute anything” mode actually covered the two plain drop statements sitting in that plan, and against a system real people depend on, that’s the gap where the risk hid: not in the step I added on purpose, but in the one line of that step I never tested before I trusted it.