73 Percent

73 percent. That was the number that stopped me, because a missing file in a five-minute update should not have had anything to do with a machine working that hard.

The immediate problem looked simple. An update was being sent to a live website. The update package was opened, and a check ran right after it to make sure a particular file, default.asp, had arrived. The tool reported that the package had opened successfully. Then, on the very next line, it said the file was missing and stopped the update.

It felt like being told that a grocery bag had been unpacked, then hearing that the carton of milk inside it did not exist. The file was in the package. The unpacking step said it had finished. Yet the next step could not see the file that was supposed to be right there.

I had an old note that said the deployment tool was broken and that a temporary patch was already in place. It sounded useful. Instead, it sent me in the wrong direction for two whole work sessions. I treated the update process as the problem and spent most of an evening trying to work around it.

That first idea did not work because it was not true. The update process was not broken. The note was stale, but it had the confidence of something I had written down for myself. Once I accepted it, every failure looked like more proof that the tool was to blame. The cost was not a bill. It was time, attention, and the slow irritation of trying the wrong door again and again.

The better question was much smaller: why would a file that had just been written not be visible a moment later?

The answer had a scary name: a race condition. It just means two steps that usually happen in a dependable order briefly get out of step. Think of putting a plate into a kitchen cupboard while someone else turns around to check whether it is there. Usually the plate is already on the shelf. On a bad day, they look during the tiny moment before it is fully in place.

Here, the machine was being asked to do too much at once. The file-opening step could say it was finished before the machine had made the new file fully visible to the next check. On a calm machine, that tiny delay never showed up. I could repeat the update by hand later and everything looked normal. Under pressure, though, the check sometimes got there first and declared the file missing.

The pressure came from a separate problem happening at the same time. The site had become slow because a large group of automated visitors were tapping the same page. More than 55 separate sources were involved, but each one only showed up once or twice. That matters because a simple report that lists the busiest individual visitor will not make a crowd like that stand out. It is like trying to find the cause of a packed shop by looking only for the one customer carrying the most bags.

Once a traffic block was put in place, the numbers changed almost at once. The machine’s workload fell from 73 percent to 11 percent. The number of active connections dropped from roughly 1,000 to 128. The site calmed down, and the updates started working again without any change to the update tool.

That was the part I had missed for most of the evening. The slow site and the missing file were not two unrelated headaches. They were the same problem showing up in two places. The flood of automated visits left so little room for ordinary work that a routine file check became unreliable.

There is a sensible improvement still to make. Between opening the update package and checking for the file, the process should pause briefly and try again a few times. That gives the machine a fair chance to finish the job before it is judged. But that improvement is still a proposal, not a completed fix. What made the updates work again that night was reducing the traffic pressure. It would be too easy to tell a cleaner story and pretend the update process had already been strengthened. It had not.

I also learned not to trust a green success message by itself. The update tool had reported success even when the right file was not on the machine. The fix was a marker placed inside the new version of the file: present after an update, the update is real; missing, the green checkmark was only ever a claim.

The number that started this was 73 percent CPU, pushed there by 55 sources nobody was watching as a group. The number that ended it was 11 percent, after one traffic block, with the update tool never touched. Nothing about the deploy process was ever broken. The machine just never had the room to prove it.