Six endpoints, one truth, and 3,673 rows that didn’t agree
The number that stuck with me was 8. Out of 3,673 session rows scattered across six endpoints that each answered “what’s running right now” differently, only 8 rows were the ones that mattered: the live agent sessions actually in flight. The data underneath was never wrong. The problem was that nothing on the page pointed at the 8 rows directly, so every time visibility broke, the fix was to add a seventh way of looking at the same pile instead of fixing the pile.
I’d been running the dashboard hard the day before, the kind of day where you spin up a dozen subagents in a loop because a central hub is too slow to wait on, and then you have no way to tell which of the dozen are actually doing something versus stalled. That’s not a nice-to-have complaint, it’s the actual failure mode: you can’t trust parallel agents you can’t see, so you either babysit them one at a time, which defeats the point of running them in parallel, or you stop trusting the dashboard and go back to a single terminal. I’d been doing the second one for weeks.
Diagnostic path: it wasn’t the data
First instinct with a dashboard that “feels broken” is to check the query. I did. The data was fine. What wasn’t fine: a subagent recorder that had been silently dead since a code change on August 22, a full week earlier, and nobody had noticed because there wasn’t one canonical place to notice it from. There were six. Two composer components at the bottom of the page were independently reading overlapping slices of the same session table and rendering slightly different subsets, so a session could be visible in one panel and absent from another, and that discrepancy read as “flaky” rather than “there are two sources of truth here.”
None of this was new. Back in June, a design review of this same cockpit had already flagged the actual fix: build one aggregation layer, a single read-only endpoint that fans out to the existing resolvers in parallel, with an explicit requirement that a slow downstream source return a partial or error state for its own zone instead of hanging the whole page. That recommendation was sound and it was never fully built. What got built instead, over the following weeks, was targeted patches on top of the old endpoints, because each individual bug had a small, obvious, local fix, and the aggregation layer was a bigger job that kept losing to whatever was on fire that day. By August 29 there were six of those local fixes stacked up, each correct in isolation, collectively presenting six different answers to the same question.
The fix: subtract, don’t add
The change that shipped wasn’t a new endpoint. It was one canonical live-session view placed on the dashboard’s front door, replacing the six panels rather than joining them. Building it forced three things that had been avoidable as long as there were six places to patch around them:
- The dead subagent recorder got rewired into the same path the new view reads from, so a recorder outage now shows up as an empty or stale view instead of silently not existing.
- The two competing composer components at the bottom of the page got merged into one. There was no way to merge them without first deciding which of their two definitions of “active session” was correct, which is exactly the decision that had been deferred for weeks by having both.
- A blog module that had nothing to do with session state, but had been living on the dashboard since some earlier “just add it here” moment, got moved off entirely to its own page. It wasn’t part of the session-visibility problem, but it was part of the layering problem: the dashboard had grown to 4,481 pixels tall, most of it content unrelated to the thing the page was supposed to answer. Pulling it out collapsed the page to 1,256 pixels.
The alternative that lost, again, was patching the sixth endpoint to agree with the other five. That’s the option that had won every previous time, and it’s exactly why there were six of them. The only way to break that pattern was to delete panels, not add one.
Verification here mattered more than usual, because “the dashboard says it’s fixed” was the same sentence that had been wrong six times before. So the check wasn’t a green test suite, it was driving the actual deployed page: opening the live site, confirming the recorder was populating the view in real time, confirming the merged composer showed one consistent session list instead of two, confirming the blog page rendered independently at its new address.
Two gaps came out of doing the work instead of just closing the ticket. One: the canonical view has no fallback if you land on the front door without the keyed link that authenticates it, which is a real gap, not a cosmetic one, and got its own ticket rather than a follow-up comment. Two: a duplicate post surfaced on the gate board during the blog move, unrelated to session visibility but only visible because someone was finally looking at that corner of the page directly instead of through a patch.
The pattern worth keeping is not “consolidate the dashboard.” It’s that a fix which makes the local bug go away without removing the surface that produced it is a deferred version of the same bug. Six endpoints each did their assigned job correctly. The correctness of each one is what let the accumulation hide for months. The question to ask before shipping the seventh patch isn’t “does this fix the symptom,” it’s “does this reduce the number of places someone has to look to trust the answer.” If it doesn’t, it’s not a fix, it’s a new place for the next person to check.