On Aug 18 the work log showed 123 sessions across 6 repos in a single day. The same day I cut one working queue from 33 to 18 items. Elsewhere, the ticket-creation default had quietly put 34 items into a working queue that nobody chose.
Coordination is what does not scale for free. Every agent I add brings capacity, and also another party that can collide with the others over the same files, the same connector, the same vault. I saw it on Aug 17. I migrated ten infrastructure tickets out of the wrong tracker and closed the originals. The next morning 3 of them had been reopened by a session that was still running, and 9 new findings had landed in the wrong place too.
That is an old problem in distributed systems, and it is the reason locks, queues and transactions exist.
The attention it takes to keep agents from stepping on each other does not shrink on its own. It shrinks when you build something to replace the human referee. I shipped a guard that refuses tickets of the wrong kind at creation time, and the leak still continued overnight in the session that was already running. Until something better exists, the referee is one of us.
AI Skills
Use this lesson with the AI assistant you already use
Two agents both wanted to write to the same shared vault at the same time, and the only thing preventing a corrupted file was a human watching two terminal windows and deciding who went first. The actual bottleneck was never a single agent's speed; it was that every new agent added another party that might collide with all the others over the same shared resource.
Paste the prompt, share only the context needed to answer it, and treat the result as a draft for your review. Do not include confidential information or let an AI assistant make changes without your approval.
Optional: for a visual report and saved memory, run /dxdev first.
Don’t have it? Get it at dxdev.com/skills/dxdev. The prompt works without it.
dxdev LESSON ยท paste into your AI coding agent
LESSON: Agent execution capacity scales with how many you run; coordination overhead scales with how many pairs can collide
SOURCE: dxdev.com/blog/2026-08-15_coordination-overhead-multi-agent
WHAT HAPPENED: A daily research agent job finished a fourteen-minute pass, firing off 26 tool calls including search rounds, page extracts, and a final write of its results into a shared vault directory, while an unrelated coding agent session was open in parallel against the same repository with write access to the same vault. Neither agent knew the other existed; there was no lock, no queue, and no signal between them, so the only thing preventing a collision was a human watching both terminal sessions and manually holding one back until the other's write finished. The insight that followed was that execution speed scales roughly linearly with however many agents are run, since the model doesn't slow down running more instances, but coordination overhead does not scale that way: it scales with the number of pairs of agents that can touch the same shared resource at the same time, which grows far faster than the agent count itself as more agents get added. The immediate fix adopted was manual and deliberately modest, not scheduling vault-writing jobs to overlap, checking a timestamp before starting a second session, while the acknowledged longer-term fix, a write lock or job queue on the shared vault connector, was left unbuilt until a concrete failure, rather than a theoretical risk, justified it.
THE RULE: When adding a new concurrent agent to a fleet, execution throughput scales with however many you're willing to run, but coordination overhead scales with how many pairs of those agents can touch the same shared resource at once, which grows combinatorially, not linearly. A human manually watching two sessions to prevent a collision is not a coordination system, it's a placeholder that doesn't scale the way the agents do.
CHECK MY CODE, then report PASS or FAIL with file:line for each:
1. Any addition of a new concurrent agent or worker with write access to a shared resource, a vault, a repo, a queue, without checking whether its write windows could overlap with an existing agent's.
2. Any manual, human-performed coordination, watching two sessions, checking a timestamp before starting a second one, substituting for an automated lock, queue, or scheduling constraint on a shared write path.
3. Any agent fleet scaling plan that reasons only about added throughput from more agents, without separately accounting for the coordination overhead each new agent adds by potentially colliding with existing ones.
THEN PRINT: a table (check, PASS/FAIL, evidence, fix) + a verdict (applies / partially / OUT_OF_SCOPE / no) + the single most important next action.