The browser had been dead for 39 hours, and it was still asking our staff toolbar who it was, over and over, with a login that had expired long before.
We found it on a Monday evening, after the error-rate alert had been firing all day. The alert was right. My first read of it was wrong.
The morning read
At 7:58 AM I reviewed the error log spike. The rows were mostly expired-session errors, and we had shipped new expired-session logging a few days earlier. We had also done a lot of overnight testing that could plausibly throw the same errors. Put those two together and the story wrote itself: the new logging was surfacing noise that had always been there, our own test sessions added to it, and no customer-facing behavior had regressed. I checked that last part and it was true.
I wrote it up that way and did one useful thing. I shipped a change to staging so those log rows carried IP, URL and browser. Before that, a row said a session had expired and nothing else, which is why I could only guess at the source.
That guess cost us the rest of the day. The spike I had explained away was still running, and nobody was looking at it because I had closed the question. The alert kept firing and I kept treating it as known noise. The “no customer regression” finding was accurate and gave me false comfort, because it answered a different question from the one the alert was asking.
What the enriched rows showed
By the time I picked it back up at 5:27 PM, the new fields were in production logs. Grouped by IP, URL and browser, the flood was not a spread of test sessions. Nearly all the volume came from one address, hitting the staff toolbar’s status call, with a browser string that identified a headless browser. Our overnight testing was a small slice of the total.
The mechanism:
- An agent had launched a headless browser earlier in the week and logged into the staff side to check something.
- The agent’s own session ended. The browser process did not. Nothing owned it anymore.
- The staff login expired on the server side. The page inside the orphaned browser was still open, and the toolbar on that page polls the server on a timer.
- Every poll came back as an expired session. The toolbar’s client code had no backoff, so it retried at full rate forever and wrote a log row each time.
The 39 hours is the gap between the agent session dying and the moment we killed the process. For all of that time the browser was doing exactly what it was built to do, and every failure it hit was logged faithfully by the logging I had just concluded was harmless.
Killing the process took one command. The rest of the evening was cleanup. We deleted the flood rows from the error log so the real signal was readable again. We also shipped a hotfix with two changes. The toolbar now backs off when it gets an expired-session response instead of retrying at full rate, and the error row records who and where, so the next one of these identifies itself on the first look instead of the second.
The gap the sweep does not cover
We already run a nightly sweep at 4:20 AM. The night before, its resync step reported 4 active ticket sessions, 3,950 zombie thought folders and 4 stuck-active manifests. It is thorough about records. It finds sessions that look alive on disk and were abandoned, and it cleans up the bookkeeping.
It never looks at processes. A session can be marked closed, its folder swept and its manifest resolved, while the browser that session launched keeps running. The records said everything had been cleaned up. The machine said otherwise.
We filed a ticket for a reaper that matches running agent browsers against live sessions and kills the ones with no owner. It is not built yet, so right now the fix is the toolbar backoff and the better log rows. That limits the damage of the next orphan but does not prevent one.
Two causes that never added up
I should not have accepted “expected noise” without asking who was making it. The enriched rows were on staging that morning, and once they reached production I could have grouped by IP straight away. Instead I explained a spike with two plausible causes and never checked whether they added up to the volume. They did not. A handful of overnight test sessions cannot produce a flood that runs for an entire workday.
Any agent that spawns a browser or a long-lived process needs an owner recorded at launch, and something that checks the owner is still alive. Otherwise the process outlives the session that made it, and you find it in your error graph.