The alert fired on a rule we’d never named: 200 events in 24 hours through the verified-bot skip lane, and not one of them was the thing we’d come to document. We opened the ticket expecting to write down a decision. What we found was that the decision had already been made for us, quietly, in a managed rule nobody had read since May.

The ticket asked for a bot-policy writeup: which crawlers get the Cloudflare Skip lane, which get walled off, and specifically what to do about AI crawlers. Simple enough. Our custom ruleset has eight rules. Rules 1 through 3 are the skip lane: our own prod-box egress IP, Cloudflare-verified bots plus calendar/RSS importers, and a signature-validated payment webhook. Rules 4 through 8 challenge or block everything else, by file-extension probe, empty user-agent, geography, ASN, and a truncated-Chrome fingerprint.

The first pass at the doc got the AI-crawler question wrong, and we shipped that wrong answer into a comment before catching it. The assumption was that GPTBot and ClaudeBot, both on Cloudflare’s verified-bot list, would match rule 2 and skip the WAF entirely, same as Googlebot. Written up as “currently implicitly allowed.” Clean story, wrong mechanism. Cloudflare’s managed bot rule sits in front of that logic and blocks unverified crawler traffic on its own criteria, independent of our custom ruleset. It had been blocking roughly 2,100 requests a day. Nobody had noticed, because nothing downstream needed that traffic and no alert was tied to it.

What was actually in the 2,100

Before writing the correction, we ran a trace against the skip rule itself to see who was actually getting through, since that’s the one place a false assumption would show up as real events rather than a doc paragraph.

ops-cli cf trace --rule verified-bot-skip --hours 24 --all-actions

200 events in the window, the honest traffic: iOS calendar syncs hitting /data/schedule.ics for youth football and lacrosse leagues, bingbot on a team roster page, an Android UA on a CSS asset. All legitimate, all skipping the WAF correctly through the verified-bot match.

That’s the allowed side. The blocked side, the 2,100/day sitting behind the managed rule, was mostly credential scanners running with an AI crawler’s user-agent string. Same UA header GPTBot presents, none of the verified signals that would let it through. The managed rule doesn’t care what a UA string claims; it checks whether the request actually resolves to the crawler it claims to be. Spoofed traffic fails that check and gets blocked at the edge, before it reaches rule 2, before it reaches our WAF logic at all. There was a minor wrinkle in the same bucket: some user-initiated ChatGPT fetches (a person pasting a URL into a chat and asking for a summary) got caught by the same rule. Volume there was negligible next to the scanner traffic.

So the actual state, once traced instead of assumed: legitimate AI crawlers were already blocked by a managed rule operating above our custom ruleset, and most of what that rule stopped every day wasn’t AI traffic at all. It was abuse traffic borrowing the UA string because it’s currently the one nobody blocks by default.

The decision that wasn’t a decision

The live question was whether to add an explicit block-or-challenge rule for AI crawlers, positioned above rule 2 so it would win over the verified-bot skip. Editing prod. The case against it: the site is public youth-sports schedules, there’s no content worth defending from a training crawler, and our own change log already had three false-positive incidents (two geo rules, one iCloud Private Relay case) from edits that looked safe going in. A new rule on a Saturday afternoon, on a site with real users checking game times, was the riskier move by a wide margin against a threat that, per the trace, the managed rule was already handling.

We closed it as documentation only. Two files: a bot-policy section describing the skip-lane split, the AI-crawler rationale, and a revisit trigger if the managed-rule numbers ever move; and a true-up of the phase-tracking doc, which had listed Super Bot Fight Mode and the custom WAF rules as still TODO when both had been live since May. No prod touch.

One more thing came out of chasing down the ticket’s other ask, a Skip rule for social-preview crawlers like Twitterbot and Slackbot. A live trace showed facebookexternalhit already hitting rule 2 and skipping the WAF, so Facebook’s preview fetches worked without any dedicated rule. But the doc that was supposed to be our canonical reference for all of this didn’t have a Facebook line at all, and it listed three of our challenge-lane rules as outright blocks. The rule that runs it had never had its own entry.

The whole ticket started as “write down the decision we made.” What it actually needed was a trace against the live edge before writing anything down, because the decision on the page and the decision the config was enforcing were two different things, and only one of them was true.