Four minutes was enough to turn a read-only agent session into a production incident. An autonomous session ran an unbounded scan against our shared production database, and the live site stalled before the session had an answer.
At 3:41 that afternoon, I started from the site-error alerts. The visible symptom was generic application failure. The actual failure was lower in the stack: one open-ended database read had created a database wedge. The session had asked a question with no hard boundary around the work required to answer it.
That distinction matters. Read access is not harmless access. A query can leave every row intact and still consume enough locking or I/O capacity to take a live service with it.
We had made production data available because agents need real state to investigate real problems. I still think that is the right direction. But I had treated database access as a permission question. The incident made it clear that it is also an execution-budget question.
Within 91 minutes, we had shipped and live-verified three controls: a machine-level block on unbounded scans, a bounded query wrapper with timeouts, and a monitoring change that treats database stalls as an alertable operational failure.
The alert was the start of the diagnosis
The first lesson was not that an agent had made a bad choice. It was that our monitoring had described the symptom too loosely. We saw error alerts from the live site. That told us that requests were failing, but not why they were failing.
Tracing the incident led to the shared production database and then to the unbounded scan from the other autonomous session. The chain was short once we had the right layer in view:
| What we observed | What it first told us | What it actually was |
|---|---|---|
| Site-error alerts | Requests were failing | The database was wedged under an unbounded scan |
| No rollback target | It was not obviously a release problem | The failure came from a live query path |
| A read-only session | No data should have been changing | A read can still saturate the resources serving live requests |
This is why I do not want autonomous access rules to stop at read versus write. That split protects data integrity. It does not protect availability.
We moved the rule into the access path
A stronger prompt was not the fix. It would make the same failure depend on an agent remembering a policy under every new task. Manual query review would remove the autonomous part of an autonomous session. Monitoring alone would only tell us the site was already in trouble.
The control had to run before the database did.
The first layer is a machine-level block on unbounded scans. The important property is where it lives. The agent cannot decide, in the moment, that its own investigation deserves an exception. If the request has the shape of an unbounded scan, it is stopped before it reaches the shared production database.
That gate also caught its own build-time test payload. I like that kind of verification because it proves the system is intercepting the dangerous path, not just documenting an intention to do so.
The second layer is the bounded query wrapper. The wrapper gives production reads a route with explicit limits and timeouts instead of handing an agent an open-ended connection. A query that cannot finish within its budget does not get to keep consuming the budget of the live service.
The block and the wrapper overlap on purpose. A block handles the class of request we now know is unacceptable. The wrapper limits the work of requests that are otherwise allowed. One rule catches a known bad shape. The other constrains the normal path. I do not want a production safety model that depends on either one being perfect.
Database stalls now count as incidents
The third layer was monitoring. We upgraded the site watcher to sample blocked and long-running SQL every minute. Lock and I/O incidents now produce capture and alerting, rather than being inferred later from generic application errors.
That change is not glamorous, but it closes the loop. CPU spikes already receive operational attention because they are visible and legible. A database wedge needs the same treatment. If the database is blocked or a query is running long enough to threaten the site, that is an incident signal in its own right.
The incident response took 91 minutes from the initial trace to shipped, live-verified guardrails. The follow-on port for other operators was queued as separate work. I did not call the job done because a single machine had a new rule. The useful unit is the access pattern, not the one host where we first learned the lesson.
I still let agents query production data. I do not let an agent turn an investigative question into an unlimited claim on the database that serves live traffic. Those are different capabilities. Treating them as the same permission is how a four-minute outage becomes an acceptable risk instead of a design flaw.