The session that read 190 million tokens

One orchestration session that week ran 644 prompts and pulled 190,414,059 tokens off cache before it closed. Output alone was 800,620 tokens. It wasn’t an outlier. It was the normal shape of a hub session, the top-level process that reads a ticket, decides what work it needs, and dispatches subagents to do it: codex exec for a code review, claude -p for a test run, manus for a second opinion. The hub doesn’t do the work itself so much as decide how much work gets done.

The subagents are cheap to reason about individually. Each dispatch shows up in the log with its own tier and its own cost: codex exec (medium, 14003t, 78.38s), manus standard (40cr, 250.16s), claude -p (sonnet/medium, 5429t, 81.97s). So when the token bill started climbing, that’s where I went looking, at the dispatch sites, one call at a time.

Where I went first, and why it didn’t work

I spent an evening going through the review dispatch call sites, the ones wired to a single ticket’s codex review and its manus review, and pinning each one to a cheaper tier. Codex reviews that had been running at high got capped to medium. Manus calls got pinned to standard instead of whatever tier a session felt like reaching for.

It didn’t hold. The next hub session that spun up, running on fable/high for its own reasoning, opened a new ticket, decided it needed a fresh round of review, and dispatched a new batch of subagents through a call site I hadn’t touched yet. My caps only applied to the sites I’d already visited. The hub wasn’t picking a model per dispatch and staying inside some budget, it was deciding, fresh, every time, how much reasoning the problem in front of it deserved, and an opus or fable hub tends to decide that expensive reasoning is warranted more often than a sonnet hub does. Patching call sites was chasing a number that could always open a new site I hadn’t found yet.

Four independent answers, same night

At 11:00 PM on the 17th I ran the same question past four different agents in parallel, not sequentially, so none of them saw the others’ output: fable, codex, opus, and manus, each asked what the single highest-leverage change to token spend would be.

Fable (1,779 tokens, 30 seconds): “cap the hub session’s model, not the subagents’.”

Opus (2,674 tokens, 43 seconds): “stop letting Opus/Fable drive sessions.”

Codex (16,036 tokens, 38 seconds): enforce a “Sonnet root, bounded escalation” policy, on the grounds that the largest avoidable bucket was expensive top-level work, not the leaves.

Manus went straight to the base rate: 2,415 subagent dispatches logged, and the real lever wasn’t the cost of any one dispatch, it was that the hub’s model choice set how many of those dispatches got made and at what tier, times 2,415.

Four agents, four different vendors, none of them coordinating, and they converged on the same root cause in the same minute. That’s not proof by itself, but it’s a strong tiebreaker against the approach I’d already spent an evening on.

Capping the root instead of the leaves

The fix that came out of that session was structural, not another call site patch: default the hub session itself to sonnet. Opus or fable at the root now requires an explicit choice, not a default a session falls into because a ticket looked hard. The subagents keep their own per-dispatch tiers exactly as before, codex exec at medium for routine review, manus standard for a second read, nothing there changes.

What changes is who’s allowed to decide, by default, that a problem deserves the expensive path. Every dispatch traces back through a hub decision: how many subagents to spawn, whether to escalate a review, whether to retry. Constrain the model making that decision once, at the root, and every downstream decision it makes is constrained with it. Guarding 2,415 individual dispatch points means the 2,416th one, at some call site I haven’t written yet, is uncapped by default. Guarding the one session that decides to make all 2,415 calls means it isn’t.

The mistake wasn’t the diagnosis, the token counts were real and the spend was real. The mistake was where I aimed the fix: at every place cost showed up, instead of at the one place it originated.