My prod box had been pegged near 91% CPU for half an hour, and the offending traffic was spread across four dominant /16 netblocks with no two requests coming from the same IP. Blocking those IPs one at a time is whack-a-mole, and the mole wins. The thing that actually killed it in fifteen seconds was not a smarter rate limiter. It was a WHOIS lookup that put the same German maintainer handle behind five separate allocations, which meant I could stop thinking in IPs and start blocking the operator.
This is the playbook. It works because scrapers rotate IPs constantly but almost never change who they rent those IPs from.
First, confirm it is actually a swarm
Before you block anything, you have to be sure you are looking at a coordinated scraper and not just a busy afternoon. High request volume on its own proves nothing. I use a three-signal triangle, and I want all three before I call it.
Signal one: URL fan-out. Real users hit a mix of pages and pull static assets. This traffic was 100% one endpoint, /teams/?u=<TEAM>&s=<sport>, with the username slot enumerating through values like BULLDOGS14, BULLDOGS31, BULLDOGS42, BULLDOGS43. Sequential enumeration of one parameter with zero requests for CSS, images, or JS is not a browser. It is a crawler walking a list.
Signal two: netblock clustering. The source IPs clustered into a handful of /16 ranges: <cluster-1>/16, <cluster-2>/16, <cluster-3>/16, <cluster-4>/16. This is the signal that gets missed if you group by individual IP. Group by /16 first. Distributed swarms spread themselves across dozens of sibling addresses inside one block precisely so a per-IP view never sees the pattern.
Signal three: header homogeneity. Every request advertised Chrome 134 or 135, but none of them sent Sec-CH-UA client hints (which a real Chrome on a secure origin normally sends), and none sent a Referer. A whole fleet claiming to be a modern Chrome while uniformly missing the headers a modern Chrome sends is spoofed, full stop.
Three signals, all positive. This is a swarm. Now figure out who owns it.
WHOIS the netblock, not the IP
Here is the move. Take the base of each cluster and WHOIS the allocation, not the individual address. You are looking past the IP at the registrant and, more importantly, the maintainer handle.
Run it on the first cluster and the same registrant falls out:
- maintainer: a single registry maintainer handle
- org: a single RIPE organization id
- a commercial proxy reseller, registered in Germany
Run it on the second cluster. Same maintainer. Third, fourth, fifth. Same maintainer every time. What looked like a handful of unrelated /16 ranges scattered across the address space resolved to one operator holding five separate /19 allocations:
<range-1>/19<range-2>/19<range-3>/19<range-4>/19<range-5>/19
Each individual offending IP I’d seen sat somewhere inside one of these five RIPE allocations. The pivot didn’t surface a brand-new range, it collapsed four messy /16 clusters plus a smaller fifth into five clean operator-owned blocks, so I could stop blocking the specific addresses I’d logged and block the whole inventory the operator can draw from. That is the value of operator-level thinking. The next IP the swarm rotates to is still inside one of these five allocations, so it’s already covered.
The ARIN-to-RIPE gotcha
One trap worth naming. A 65.x or 209.x address looks American and your instinct is to query ARIN. ARIN will hand you back a referral, because the block was transferred to and is now maintained under RIPE. If you stop at the ARIN response you get a vague legacy record and miss the maintainer handle entirely. When the registry you queried points at another registry, follow the pointer and re-query the right one. The German LIR identity only shows up in the RIPE record.
Block the allocations as units
The mitigation is one firewall rule per RIPE-allocated range, scoped to the whole allocation, not per IP. On a Windows box that is New-NetFirewallRule:
New-NetFirewallRule ` -DisplayName "Block-ProxyOperator-<base>-2026-05-26" ` -Direction Inbound ` -Action Block ` -RemoteAddress <hosting-range>The naming convention is doing real work here. Block-ProxyOperator-<base>-<date> tells the next person (or the next me, six weeks from now) exactly what this rule is, who it targets, and when it went in, without opening anything. The rule is self-documenting by its own name.
One detail that matters more than it sounds: write this as a .ps1, copy it to the box, and run it with -ExecutionPolicy Bypass. Do not try to author five New-NetFirewallRule calls as an inline SSH one-liner with nested quoting. That is how you end up debugging a parser error while the site is down. The script file parses cleanly the first time. It is the fast path, not the careful one.
The result
CPU went from 91% to 13%. Connections dropped from 490 to 146. Slow requests over sixty seconds went from 22 to 1. All of that in about fifteen seconds of the rules taking effect. Roughly 40,000 addresses blocked (five /19 allocations is about 8,192 addresses each), and the legitimate-customer blast radius is effectively zero, because the entire allocation belongs to a proxy reseller. Nobody’s actual customer is browsing your sports site from inside a commercial proxy LIR’s whole /19.
The durable artifact is the maintainer handle
The block is the cheap part. The expensive part was the investigation: the three-signal triangle, the WHOIS pivots, the ARIN-to-RIPE re-query. I did not want to pay for that twice.
So the real output of the incident was not the firewall rules. It was a one-line recognition key written down where I’ll see it next time: that maintainer handle is a known commercial proxy reseller, and any range under it is fair game to block on sight. The next time traffic from that operator’s allocation shows up, the work is a thirty-second recognition, not a fresh half-hour investigation.
That is the shift. Per-IP blocking treats every wave as new because the IPs genuinely are new. Operator-level blocking treats the wave as a known quantity because the operator almost never changes. The IP is the disguise. The maintainer handle is the face under it.
A note on band-aids
Firewall blocking is a band-aid and I will not pretend otherwise. The proper fix is an edge solution that does this for you before traffic ever reaches your origin. That work was already in flight as a separate multi-week effort. But a fifteen-second block that restores a downed site right now is not a failure of architecture. It is correct triage while the real fix gets built carefully. Ship the mitigation to stop the bleeding, file the proper fix as the real work, and feed every incident’s intel into both layers. I also filed a follow-up to ingest these exact /19 ranges into the app-layer IP table, so the next time one slips the edge, the application flags it too.
Takeaway
On the next wave, start with the maintainer handle recorded for the five /19 allocations. Re-query each source range in RIPE and confirm it returns the same operator before adding a firewall rule.
Related
The recognition key that ended this incident was one German maintainer handle sitting behind five separate /19 allocations, roughly 40,000 addresses in total, and that handle is now the only thing worth checking on the next wave. Confirm the same registrant comes back on a fresh WHOIS pivot, write the five ranges into a firewall rule scoped to the allocation rather than the address, and the next fifteen-second fix is really a thirty-second lookup against a fact you already have on file.