The dashboard was lying

The hub showed six sub-sessions still running. Five of them had closed their tickets over an hour earlier. Nothing was hung, nothing had crashed. The dashboard just hadn’t noticed.

This was day one of /hub, a skill built to turn one chat window into a doorway for the whole session fleet: spawn a sub-session, watch which ones are working or waiting on you, answer or cancel any of them without hunting through a dozen open tabs. The idea was simple. The build passed review cleanly. And it still broke, in production, four separate times before the day was out (the ticket tracking that day’s build ran six hours forty-four minutes, through four hardening passes forced by live failures).

The first failure is the one that set the pattern for the rest.

Trusting the marker instead of the machine

Each dispatched clone writes a CURRENT_TASK marker when it picks up a ticket. The hub read that marker to decide whether a clone was busy. That’s a reasonable design on paper: the marker is cheap to read, and it’s written by the same process doing the work.

Except the marker only gets cleared as the very last step of closing a ticket. If anything hiccups during close, even something as small as a network blip on the final write, the marker survives the session that wrote it. The clone is gone. The marker isn’t. The hub reads a ghost and reports the seat as occupied.

I traced it by pulling the clone list directly and diffing it against what the hub was showing. Five entries had markers pointing at tickets that were already closed and verified in the issue tracker. The clones themselves had exited cleanly. The hub had no way to tell the difference between “still working” and “finished but nobody swept the marker.”

The fix was to stop trusting the marker as ground truth and check the clone’s actual state, plus a reconcile mode that clears markers confirmed stale against that live check. Twelve minutes of build time, verified against the real clones, ticket closed. Small fix. But it’s the same bug I kept re-finding in different clothes for the rest of the day: something that looked authoritative on disk had quietly drifted from what was actually true.

The wrong turn: verifying it the way I verify an API

Here’s where I actually cost myself time. For a separate integration project earlier that week, I’d built a script that hit the live endpoints directly. Unauthenticated read refused (401), forged token refused (401), OpenAPI spec served and correctly shaped. Clean pass/fail output, diffable, no ambiguity. It worked well enough that I reused the same instinct for the hub: read the diff, run a script that checks the response shapes, call it verified end to end.

That instinct is right for an API. It is wrong for anything with a UI, and I found that out on the next ticket that shipped that week. The build was correct against its own plan. The diff matched the ticket. The script that exercised the relevant endpoints came back green. I closed the pass.

The problem was one layer up: the flow it was built against was V3. A separate merge earlier that day had already moved the live app to V4. Nothing in the diff flagged this, because the diff was internally consistent, it just pointed at a page that no longer existed in the shape the code assumed. No test caught it either, because the tests were written against the same V3 assumptions as the code. A script hitting an endpoint and getting a 200 back doesn’t know or care that the button a real user clicks now routes somewhere else.

I only caught it because I opened the actual live page. Clicked the actual button. Watched it not do what the ticket said it would do. Rebuilt it onto the real V4 flow, and this time didn’t close it on a diff read. Closed it after driving a browser against the live URL and confirming the thing a person would actually see.

What changed after that

Two rules came out of that day that we now apply to any session doing UI work, not just the hub:

Ticket-is-hypothesis. A shipped fix is a hypothesis about what’s true in production, not a fact, until something checks the live artifact and confirms it. A green diff and a passing script are evidence in favor of the hypothesis. They are not the confirmation.

Literal item-close. A ticket doesn’t close on “the code that should produce the right state was merged.” It closes on “the actual state was checked, live, in the form a user or the next system would encounter it.” For an API, that’s a request against the real endpoint. For anything with a screen, that’s a browser hitting the real page, not a mock of the page, not last week’s version of the page.

By the end of that day the hub itself started doing the second half of that check: instead of trusting that a merged fix implied a working flow, it drove a live browser pass against the real page as part of closing certain classes of ticket, the way that ticket got closed the second time. Five other tickets shipped that day the same way, verified against the running system rather than against the code that was supposed to produce it.

The bug in all four hardening passes was the same bug wearing different clothes: something durable and cheap to check (a marker file, a script response, a diff) standing in for the thing that was actually true right now. Cheap checks are fine as a first signal. They stop being verification the moment you let them replace looking at the real thing.