The Test That Never Ran but Passed Anyway
The full test suite takes 497 seconds. For as long as anyone had written it down, the number in our notes was 26 minutes, and nobody had ever timed it.
The build log
Build log, architecture patterns, and observations from running autonomous AI systems in production.
The full test suite takes 497 seconds. For as long as anyone had written it down, the number in our notes was 26 minutes, and nobody had ever timed it.
Blue outlines stuck on a calendar in one customer's Firefox. The first hypothesis pointed at her machine. The outline came from a class our own JavaScript toggled on hover, and moving it to CSS ended the reports.
A hook logged work whenever a ticket key showed up in text. A second reviewer caught the flaw. It went live a day later, with receipts.
At 6:58 AM on the sixth day, I found the nightly blog drafter had been dead since the previous week.
At 12:45 in the morning, the first review came back in three minutes: the funnel has no top.
47 posts went live on verdicts that said the post needed work. I found that out by chasing a failed blog cron, which took a 41-hour session to run down.
Five nights running, the nightly blog drafter produced zero posts. Five nights running, the only signal we had was a CI triage bot repeating the same line every 45 minutes: publishPost returned ok:...
The Railway build log read:
My agent kept telling a teammate's agent to pull a fix it already had. The real problem was that its posts had no mention, so nothing woke to read them.
A client's plan offer expired after 23 hours because it only existed as an unpaid Stripe subscription. The fix I had skipped rested on an objection nobody had checked.
The real failures and fixes from building AI systems, one practical lesson per post. Get the next one in your inbox.