Our router couldn’t see one of our two paid Max 20x accounts for 809 minutes, about 13.5 hours, and none of its default output said so.

We spread agent work across those two accounts, and their weekly windows are staggered. One resets Thursdays at 08:00 UTC and the other Sundays at 10:00 UTC. A router decides which account each launch lands on, using pacing samples: used percent per window, reset time, and a projected exhaustion date. When one account is on track to leave quota unused, work should spill onto it.

The wrong question

On Sunday we expected the Thursday account to have reset with all its credits back. It hadn’t. Live output said 69% of the weekly window used, next reset Thursday 08:00 UTC, and the router projected it would run dry Tuesday around 20:20 UTC. That projection is a slope, not a measurement.

So the question we started with was why that account hadn’t reset. It was the wrong question. The reset we remembered had happened on the other account, that morning at 10:00 UTC. We couldn’t see it happen, because the router had been blind to that account since the night before, when its last sample read 100%.

What the blindness looked like

The plain pacing command prints a verdict per account. Nothing in it announced a problem. The sample-only subcommand prints per-account errors, and that is where it showed up:

hq pacing sample
account B: OAuth token has expired and has not been refreshed yet

The routing command was more candid:

hq seat-route --venture <name>
account B pacing is 809m old, so pacing cannot be trusted

So the router did the right thing with stale data and refused to use it. The cost was that a freshly reset weekly window sat unrouted on one account while the other burned toward early exhaustion. Nothing spilled over, because as far as the router knew, one of its two accounts had no data.

The mechanism is small. The sampler reads the access token and never refreshes it. Only a real claude call under that account’s config directory refreshes it. An account that sits idle past its access-token life drops out of routing. And the account most likely to sit idle is the one with headroom, which is the one you most want to send work to.

We checked the long-lived credential separately. The seat-watch command reports days until each refresh token itself expires, and both were healthy at 25 days and 22 days. The refresh token was fine. Nothing had used it to mint a new access token.

The smoke test that failed once

To confirm the account was spawnable, we ran a minimal call under each one:

hq claude --seat B --model haiku --effort low "Reply with exactly the word OK."

Account A returned OK on the first try. Account B failed with:

Failed to refresh OAuth token: another Claude Code process is refreshing it or exited mid-refresh

We retried, and it returned OK. That retry was the fix, because the call refreshed the token. The next pacing sample picked account B up: 0% weekly used, resetting Sunday 2026-09-27 at 10:00 UTC. We don’t yet know whether the first-attempt refresh failure recurs, and we haven’t checked.

What we haven’t built

Nothing here is fixed yet. The plan is untested:

  1. Have the sampler, or a headless cron, fire a tiny call under any account whose token has expired, so the router can’t lose an idle account.
  2. Send bulk work to the account with the most headroom relative to time left before its reset. The router already has this logic. It only needs fresh data from both.
  3. Make staleness a loud state in the default output instead of a line you get by asking the right subcommand.

There is also a decision we haven’t made: whether work tied to one funder is allowed to spill onto the other account at all, since account choice matters for attribution.

One number we can’t yet use: over the last 7 days, both accounts together ran 2,049 sessions and 157,718 model turns, with 20.3B cache-read tokens. Subagent runs were 1,422 of those sessions and about 64% of the cache-read. Cache-read isn’t the same as quota cost, so how much of that 69% subagents explain is unverified. These figures also come from the usage endpoint the sampler polls, and we haven’t cross-checked them against the vendor’s own usage page.

A monitor that stops seeing something is reporting a fact about itself, and ours already had the sentence for it. It was just written where you’d only read it by asking.