The HVNABA scrape passed every signal in my /bot-attack skill. Volume was moderate, response codes were clean, and the user-agent strings looked plausible. The IPs were residential addresses across North America, not the Hetzner or DigitalOcean ranges I had learned to recognize on sight. My detector was not broken. It was calibrated for a different threat shape.
My /bot-attack skill is a Claude Code session-starter that pulls the last N minutes of IIS logs, groups by IP and path, checks against a block list, and produces a triage bundle. It works reliably against datacenter scrapers. A single /24 shows up in 800 requests inside a minute, the pattern clusters, and I can fire off Windows Firewall rules in 20 minutes. That was the 3xK Tech GmbH wave from the prior week. This was not that.
HVNABA ran 43 unique IPs in the first peak window, with no two sharing a /24. Each IP individually looked like a slow human session: one or two requests, reasonable gaps between them, no repetition. Together they were hitting /teams/?u=<TEAM> across my entire roster in alphabetical sequence, one request per IP, cycling through teams in order. The pattern was only visible when I grouped by path template and sorted by request time across the full 11-minute window. Per-IP, every one looked fine.
I had been modeling scrape threats by source density, and residential-proxy swarms are designed specifically to defeat that model.
How the first passes landed wrong
I ran initial triage with Claude handling the analysis layer. The first pass flagged volume and a suspicious IP from a shared hosting block. The second pass caught the path-cycling pattern but classified it as a poorly-behaved search crawler, not a coordinated scrape. The third pass acknowledged the residential IP distribution and suggested bot-score throttling as the mitigation.
Three passes in and we were still proposing heuristics on top of heuristics. Each new heuristic is calibrated to a prior incident. They accumulate. They start contradicting each other. Eventually a sufficiently novel scrape walks through the intersection of all of them because no single rule fires with enough confidence to block.
The asset-mix observation broke the loop. I looked at what HVNABA was actually requesting: roster pages only, zero static assets, zero CSS, zero images. A real browser session pulls dozens of static files. A session that fetches 12 HTML pages and zero images is not a browser, and that signal requires no threshold, no prior data, and no tuning against historical incidents.
Once I had the asset-mix read, the classification was clear. But I had no way to encode it cleanly, and no way to verify that any future reimplementation would catch the same pattern. I stopped writing mitigation rules and wrote Verdict_Schema.md instead.
Writing the contract before the detection code
Verdict_Schema.md defines the structure a detection verdict must produce: verdict type (swarm, crawler, human, ambiguous), confidence between 0.0 and 1.0, the primary evidence signal that drove the verdict, a secondary confirmation signal, and the recommended action. It contains no detection logic. It says nothing specific about HVNABA. It is the output shape that any implementation must match.
The reason I wrote this before touching any code is that two implementations were coming. I planned to run both Codex and Manus against the same detection problem and compare their approaches. Without a shared schema, “Codex says swarm, Manus says aggressive-crawler” is an anecdote. With Verdict_Schema.md, both produce comparable JSON and I can diff them on the same fixture.
The first HVNABA peak window became that fixture: 43 IPs, 612 requests over 11 minutes, expected verdict swarm, confidence >= 0.9, primary signal asset_mix. Any implementation that scores this slice below 0.9 on swarm fails the contract. The fixture is a labeled example of a real event, not a synthetic test case, and it stays real no matter how many times the detection code gets rewritten.
Phase 1 is not done
The full HVNABA session log has 12 peak windows. I calibrated the fixture against the first one. The remaining 11 are unsorted. Some may show different behavior as the scrape evolved over the session. Some may force a second verdict class, something between crawler and swarm that the current schema does not name yet. I cannot know until I label them.
The cross-implementation work is also unfinished. Codex has a detection draft. Manus has not seen the schema yet. Getting both implementations to produce passing verdicts on the same fixture, built from real traffic, is the actual validation, and it is slower than writing either implementation individually.
This is the state Phase 1 is in: one fixture passing, 11 windows unreviewed, two implementations without a shared test run. I am writing it down because “Phase 1 complete” in solo tooling usually means “I wrote the code.” Here it means something more specific than that.
What the schema actually protects
Every heuristic in /bot-attack encodes a threat pattern from an incident I remember. I cannot explain half of them without reading the source. They accreted over two years of scrape waves, each one plugging a gap I discovered after it had already cost me something.
The fixture from the HVNABA peak window sits differently. It is 43 IPs, 612 requests, labeled, versioned, and required to pass. I cannot ship a new Codex or Manus draft without running it against that specific slice of real traffic. I cannot accidentally regress the residential-proxy detection without a visible test failure. The contract gives the detection code a way to fail loudly instead of drifting quietly into a gap I will not notice until the next morning’s CPU chart.
Verdict_Schema.md is the line where HVNABA stops being a story I tell about a scrape wave and becomes an obligation that the next implementation either passes or fails.