At 10:29 one morning I wrote the status line for a ticket, and it had to say two things that didn’t fit together. One half of the work was finished, tested and pushed. The other half was marked done in our shared record, but it was missing 14 fixes that we had already verified.

What the ticket was for

We run AI agents overnight for a lot of small jobs. They answer tickets, review code, and check each other’s work. Earlier that same night I had found that they were burning through an account with no quota left and had no idea. One of them retried a maxed-out login for hours because the “you’re out” message arrived in a place the code never looked. It kept trying and never registered a failure.

The ticket that morning was the fix. When one login runs dry, an unattended job should slide over to another one. It should never fall back to a pay-per-use key that bills by the call. A flat subscription that hits its limit costs you a delay. A pay-per-use key costs you money, and nobody is awake at 3 AM to notice.

How 14 fixes went missing

Several sessions had been working on that one ticket in parallel that night. Each one was an AI agent with its own workspace, and none of them knew about the others. The main session had been at it for over eight hours. Over that time it built the full version: walls that keep trial runs away from the live setup, a fix for one of our tools sneaking past the spending check, a fix for the system claiming we were rate-limited when we weren’t, and 11 other corrections that a second agent had attacked and failed to break.

Then the machine was about to reboot. We have a habit for that. Before a restart, a sweep goes through and saves everything uncommitted, so a reboot can’t eat hours of work. It’s a sensible habit, and it’s what did the damage.

The sweep saved an older, incomplete copy of the seat-failover work to the shared main branch, which is the official version everyone builds from. It had none of the three headline fixes and none of the 11 others. It landed as a normal-looking save with a normal-looking label, and nothing about it said “unfinished.”

The finished version was never lost. It sat uncommitted in a separate working folder, untouched, along with a new shared helper that closed one more gap in stripping pay-per-use keys out of the agents’ environment. But the official record said one thing and the folder said another. Anyone, human or agent, who picked up the ticket from the record would have started from the weaker copy and believed it was current.

To make it worse, a competing fix for part of the same problem had landed on its own in a different project’s main branch that night. Nobody had asked for it. Two sessions had solved overlapping pieces without either one knowing the other existed.

What we got wrong

We built the pre-reboot sweep to protect us from losing work, and I had tested it against exactly that failure. I never asked what it would do when the work in the folder was newer than what any session had finished writing, or when two sessions were writing near the same files. It did what I told it: save what it finds. It couldn’t know that what it found was a half-finished snapshot of somebody else’s eight-hour job.

A save button that saves whatever is on the desk isn’t a backup of the right version. The cost was small this time only because the complete copy happened to sit in a folder the sweep didn’t touch. If it had been in the same place as the stale one, the good version would have been overwritten, and 14 verified fixes would have become 14 things to rediscover from memory and notes.

The rest of the damage came in smaller pieces. The record was false for a while. And the ticket, which is about not spending money we don’t intend to spend, couldn’t be closed, because half of it was sitting in a folder and the shared version was wrong.

Where it stands

The half of the ticket in the other project is done. It’s pushed, the tests pass, and the change that strips pay-per-use keys from the bot’s environment is live. The half in our main project is not done. The complete copy is safe and uncommitted, and I’m writing this before I’ve replaced the stale commit with it. I’d rather tell you that than dress it up.

Here’s what we’re changing first. A session that owns a ticket now has to say so somewhere the others can see, before it starts. The reboot sweep gets a rule to skip any folder that belongs to a session still working, and to flag it instead of saving it. If two sessions touch the same ticket, that gets flagged too, instead of leaving me to find it at 10:29 in the morning by reading a log.

If you run more than one helper on the same job, whether those helpers are people, contractors or AI agents, ask who is allowed to save the shared version. If the answer is “whoever gets there first,” you have the same problem we did. Ours cost us a morning and an untrustworthy record.