The 683 Lines Before the Repair

By the time I opened the backup, a deleted site had already become a recovery involving 327 teams.

At first the problem sounded small: a customer’s site was gone, and the usual picture that brings to mind is a few pages restored from a saved copy. That was not the picture here. The site was a large web of information spread across many places, games, scores, news, rosters, statistics, and every team belonged to the same league, with records pointing back and forth to each other. Restoring it meant writing new copies into the live system customers use every day, while making sure every internal reference still pointed to the right new place. Getting that wrong once would be bad. Getting it wrong the same way 327 times could turn a recovery into a much larger cleanup.

So I did not start by changing anything live. I made the recovery walk through the whole job with its hands in its pockets first: read the backup, find all 327 teams, count everything it would need to bring back, and write down every decision. The result was a 683-line plan file. It said what the recovery would do, team by team, without doing any of it yet.

That list mattered because a message like “everything looks good” would have been useless. A list can be checked against the real result afterward, line for line. If a team was promised in the plan but missing from the recovery, the gap would show. If seven games were expected but fourteen appeared, that would show too.

That second example was not imaginary. An early test exposed a real mistake: one team had seven games in the backup, but the target ended up with fourteen. They were duplicates, copies of the same games under new labels. The first version of the recovery had no protection against being run twice, and I had not thought to add one until the plan-versus-result comparison put the mismatch in front of me. The cost of missing it would have been time at the worst possible moment, with a customer waiting and a manual cleanup in the live system hanging over the work.

The fix was not a note to remember not to run it twice. Each new game was marked with the label of the game it came from. Before adding a game, the recovery checked whether it had already brought over the one with that label. If it had, it reused the existing one instead of making a duplicate. Now a second run could be checked in seconds by counting, not by digging through a pile of nearly identical records.

There was one more design choice that kept the risky part from getting loose. The recovery had a single setting with two positions: plan or run. Plan meant read and write the list. Run meant make the changes. There was no separate switch for whether the script was allowed to write. That question was answered entirely by which position the one setting was in, so there was nothing to forget to also flip back.

The plan file proved the traversal was sane, but it could not prove the writes themselves were correct, because in plan mode nothing gets written. So before touching the real customer’s account, one full recovery ran against a disposable copy of the database made specifically so it could be thrown away. There, the recovery was allowed to actually write, and three plain checks ran against the result afterward: did the game count match the backup, did every new record carry its original label, and were there any duplicate pairs that should have been only one. Only after one account passed all three, on a copy that didn’t matter, did the recovery run against the real one.

When the real run finished, it produced a second file: 988 lines, one line per team recovered. Now there were two documents describing the same job from two sides. One said what should happen. The other said what did happen. Comparing them made “did this actually work” something you could answer by reading, not by trusting.

The last part came after the recovery had already worked, and it’s the part I was most tempted to skip because the job was done. The script was still pointed at a real account, and its setting was still on run. So the last step was resetting the target to a placeholder and the setting back to plan, so that using it live again would take two deliberate choices instead of one leftover one.

If you run something like this

You don’t need 683 lines. Most bulk changes don’t touch 327 of anything. But the same three questions apply to any job that repeats one change across many records, whether that’s a data cleanup, a bulk email, a price update, or a batch of refunds:

  1. Can it tell you what it’s going to do before it does it, as a list you could actually read, not a count you have to trust?
  2. Can it tell the difference between running for the first time and running again, or will a second run quietly double the damage?
  3. Does turning it off take one step, or does someone have to remember two separate ones?

The 327-team recovery answered yes to all three, and the second question is the one that actually caught a real mistake before it multiplied 327 times.