The Test That Passed in Fixtures, Failed in Production
The alert fired three times in a row: check:rls exiting 1 after every listed check printed pass.
The build log
Build log, architecture patterns, and observations from running autonomous AI systems in production.
The alert fired three times in a row: check:rls exiting 1 after every listed check printed pass.
524s on the main site, measured over five complete days: 6 on the 24th, 9 on the 26th, 1 the day after. The Dev Notes on the ticket had called it one account's data hitting a slow path.
Forty-three: that was the number that kicked off the post-mortem. One day, I'd closed forty-three agent windows manually, one at a time, no batch action, no shortcut.
The alert wasn't a bug report. It was our support lead, on the phone with a customer for hours, because our own screen was telling her he was wrong.
Both claude.ai session URLs bounced to claude.ai/logout?involuntary=1 before I had typed anything, and the login page carried returnTo=%2Fcode%2Fsession... so the redirect was clearly deliberate.
The report said "critical bug on the live tournament, the design dropdown is not working," and the menu started at y=96 while both of its ancestors' clip boxes ended at y=95.
The lease said held. The board said stale. Both were reading the same file, and both were right, for different definitions of "right."
At 3:41 PM the automated triage came back in 5.28 seconds, 76 tokens, thirteen cents: "cause unclear: health check failed after deploy AND rollback also failed; underlying app failure not shown in ...
One of our own published posts carried "working-queue-grew-to-65" in its slug and never once printed 65 in the body.
The new landing block showed every visitor the same pitch, including the ones who already had an account.
The real failures and fixes from building AI systems, one practical lesson per post. Get the next one in your inbox.