A teammate’s agent sat broken for two days, and the way it finally got noticed was someone squinting at Discord reactions and thinking they looked wrong. Then it spammed the channel. When I sat down with it, three separate defects were live at once, and a script already in the repo could have named every one of them in seconds.
Three defects live at once
A cleanup tool that crashed on Windows before deleting anything. The console was on the ANSI code page. One character outside it in a Discord username or a JIRA display name turned the tool into a traceback.
A race that erased why a failed model call failed. The agent knew a call had failed and had lost the reason, so it could not say anything useful about it.
A scheduled task with no repeating trigger. It was registered ONLOGON with no repetition, so once the process exited, nothing would start it again until he logged back in.
The third one is the one that hurts, because our own update tool caused it. update.py writes a restart flag, the listener sees it at the next sweep, prints “stopping so the runner brings me back on it,” and exits with code 3. On his box there was no runner to bring it back. The tool that took his agent down reported success.
The morning I lost to a wrong runbook
My first move was to unblock the fix, and I lost the morning on it. The setup docs told people they needed an administrator shell to register the scheduled task. He followed them, hit the elevation wall, and stopped. I followed the same docs, tried to walk him through elevating, and we burned hours on that path. The instruction was simply wrong. Registering a per-user task never needed elevation.
So the cost of trusting the runbook was a whole morning during which the fixes existed and could not land. I corrected the docs first. After that the update went through, and we watched a full restart cycle on his machine (stop, exit 3, come back) to see it happen instead of assuming it would.
The fix for the tool itself is a check before the request, not after:
def will_come_back() -> tuple[bool, str]: """Is there anything on this machine that starts the listener again? Asked BEFORE requesting a stop."""On Windows it wants a scheduled task with a real repetition interval. If the answer is no, update.py refuses to request the stop and says why.
The check nobody ran
None of those three bugs was the actual failure. benagent/selfcheck.py checks seven things that can each break while the bot still looks online in the member list:
- Its Discord identity.
- That JIRA sees it as its owner.
- That it can reach the model.
- That the principal’s personal Claude setup did not leak in.
- That nobody has stood it down.
- Two more that follow the same shape.
It only ever ran when a human typed python benagent/selfcheck.py. Nobody typed it for two days. Even the docs had drifted, still saying “five checks” after it grew to seven, which tells you how rarely anyone looked at it.
Making the agent watch itself
The listener already has periodic machinery (sweep_loop and idle_watchdog), so I reused it rather than adding a scheduler. The design constraints came straight from the incident:
Hourly, not every sweep. selfcheck shells out to claude -p, so each run is a real model call. Cost is a design input here. Hourly is the right order of magnitude, and it sits behind a timestamp gate inside the sweep.
Silence on pass. No post, no DM, nothing above debug in the log. Silence is the signal that the agent is healthy.
A DM to the principal on failure, never #agents. It names each failing check and its detail line. This is the owner’s problem, and nobody else in a shared channel can fix it.
Once per distinct failure. The repo already solved this. On 2026-08-04 an agent posted eleven identical warnings, which is why said_recently() and the spoken state exist, keyed by signature. I made the signature the sorted set of failing checks. The same failure stays quiet for the hour-after-hour repeats, but a new failing check changes the signature and gets heard.
One message on recovery. A latch that only ever closes is a mute. When a previously failing check passes again, the agent says so once and clears the signature, the same move as unsay(OUTAGE) for model outages.
It cannot take the listener down. The whole check is wrapped so any exception is logged and swallowed. A health check that kills the process it is checking would be a new way to fail.
To make it callable in-process I refactored the checks to return structured results (label, ok, detail) instead of only printing. The command-line entry point still prints the same [PASS] and [FAIL] lines, so anyone running it by hand sees nothing different. The plain-console guard stays, since a check that dies while reporting is worse than no check.
What changes on the next broken day
Before, a broken agent looked identical to a healthy one until someone guessed. Now the worst case is one hour of blindness plus one DM listing exactly which of the seven checks failed and why. The Windows crash, the race and the scheduled task would each have shown up as a named line within the hour.
We filed the follow-ups the same day. Two are the neighbouring gaps this exposed: nothing is read on startup, so anything posted while an agent is down is never seen, and agent-to-human messages need a length cap. The third is the heartbeat, and it is the one that changes how the whole system feels. An agent that can be silently broken for two days is one you have to supervise. An agent that tells you within the hour is one you can leave running.