Commit commit-C1BF changed one environment variable back to DB_PASS after an agent had renamed it to DB_PASSWORD and broken a database connection that had been working.
That was not the session-token fix. It was a recovery from a new failure introduced while we were already chasing another one.
The visible symptom was a 401 during token validation. The service was a Node API backed by SQL Server, with secure tokenized session forms in the request path. Two agent personas were working in the same repository. The incident runs through six commits: commit-C1BF, commit-6CC2, commit-70AC, commit-7298, commit-FBCC, and commit-89E9. From the outside, it looked like agents debugging each other in circles.
The loop had a recognizable shape. Add DEBUG logging. Declare a CRITICAL FIX. Revert a change. Add more logging. Declare another fix. Unwind the last thing that changed because the system is still broken.
I do not mean that as a complaint about agents. It is a failure mode worth naming. An agent can make locally reasonable moves for several commits while the incident gets harder to read. A second agent can inherit that altered state, treat it as ground truth, and reason correctly from the wrong starting point.
The database change was a side quest
The DB_PASS to DB_PASSWORD rename sounds harmless in a diff. One file changed, one line removed, one line added. But the service already had a working database connection. Its configuration contract was DB_PASS. Changing the application to expect DB_PASSWORD disconnected the code from the environment that was actually deployed.
The revert was right, but it illustrates why a stuck loop is deceptive. Once a debugging pass creates a second failure, the next pass must distinguish original symptoms from fallout. A 401 caused by token-path state now sits beside a database failure caused by configuration churn. The log gets louder. The causal picture gets worse.
That is the real cost of this incident, and it is not hypothetical: the response to the original bug produced a second, unrelated outage before either one was fixed. A service that started with one failing invariant ended up with two, because nothing stopped the first agent’s confident rename from shipping. The database connection had been working. It stopped working because of debugging effort aimed at something else entirely.
The discarded alternative was to treat the variable rename as cleanup and update the environment around it. That would have expanded the blast radius of a one-line change and converted a broken assumption into a new dependency. Reverting was cheaper because it restored the known-good boundary immediately.
What the agents were actually missing
The token validation problem was not a database credential problem. The commit history identifies two root causes: variable hoisting and duplicate session IDs.
That combination explains why the loop survived several plausible fixes. A token check can fail because the supplied token is wrong, expired, missing, read from the wrong field, associated with the wrong session, or evaluated against state that has not been initialized as the code expects. 401 does not tell you which of those happened. It only tells you the authorization boundary rejected the request.
The first approach treated the response as an observability problem. More DEBUG logging was a sensible option. We needed to see the request path, token lookup, and session state rather than guess. Logging lost as the main strategy because the output did not settle the identity question. It showed activity, but it did not prove that the session record was the intended one or that validation variables existed with the intended values at evaluation time.
The next option was to patch the validation path directly. That produced the kind of commit message that should make a reviewer slow down: CRITICAL FIX. A strong label is not evidence that the diagnosis is strong. The commit history later records variable hoisting as the real cause, alongside duplicate session IDs. The behavior looked like failed authentication, but the defect lived in the session state and execution order beneath it.
The other discarded option was to keep making isolated fixes until the 401 went away. That can produce a green result for the wrong reason. A duplicate session ID may point to a record that passes validation, while the data model remains ambiguous. A change elsewhere can briefly mask a bad evaluation order. Neither outcome is a root-cause fix.
The repair had to account for both conditions. The session identifier needed to resolve to one unambiguous session, and the variables used to validate the token needed to be available at the point they were evaluated. Those are separate invariants. Treating them as one generic “auth bug” is what let the loop continue.
How I spot the loop now
I do not judge a debugging run by whether it produces a confident explanation. I watch the sequence of changes.
A loop is forming when commits alternate between observation, urgent correction, and reversal without narrowing the failing invariant. DEBUG changes and reverts are not bad. They become a warning when each one changes the operating picture more than it reduces the hypothesis space.
This incident made me stop treating agent personas as independent verification merely because they have different prompts. Two agents working from the same repository can amplify the same mistake if neither is required to restate the invariant before changing code. A fresh persona is only a fresh review if it can challenge the previous diagnosis instead of inheriting it.
The guardrail we built later was a small workflow with a hard stop after three failed loops: orient, verify, fix, commit, deploy, then assess whether the last change actually narrowed the fault. The point was not to make agents less autonomous. It was to force a reset before a fourth plausible change landed on top of three unproven ones.
I built that stop after this incident, not before it, which means I let this one run the full six commits. Three of them still looked like measured progress to me at the time I read them. That is on me, not on either agent: nothing in how I was directing the work said three unproven fixes in a row is itself the signal to stop, so nothing forced a pause until I wrote that rule down afterward.
The best signal was sitting in the history all along. When an agent renames a working environment variable, reverts it, keeps logging, and then declares a different root cause, the problem is no longer only the 401. The debugging process has lost its model of the system.
That is what a stuck agent loop looks like from the outside. Not one dramatic hallucination. Six small commits that each make sense in isolation, until the commit graph shows they were all turning around the same missing invariant.