A customer’s homepage had been returning an error page for two weeks. Not to them. To anyone who typed their short address. Nobody reported it, because other ways into their site worked fine the whole time.

I found this by trying to break my own work, not by anyone reporting a problem.

A few weeks earlier I’d been closing out a run of scanner traffic hammering the platform’s edge. Vulnerability scanners probe the same short list of paths on every site they touch: WordPress admin routes, a handful of framework debug endpoints, a scattering of generic strings that only ever show up in recon tooling. The fix was a blocklist rule that matched those exact strings and dropped the request before it reached the origin server. It worked. Failed requests reaching the origin dropped hard, and I moved on to the next scanner in the queue.

One of the strings I added to that list was a short, unremarkable-looking path segment. It had shown up zero times as a legitimate 200 response across a week of logs on every site the platform hosts, which was the bar I’d set for adding anything to that rule. Zero legitimate hits looked like proof it was safe.

It wasn’t proof of the thing I needed proof of. The platform gives every customer a canonical short URL built straight from their own username, one path segment, nothing else. The string I’d blocked as a probe signature was also, by pure coincidence, one customer’s actual username. Their own front door, the exact address printed on a flyer or pasted into an email, matched the rule and got dropped at the edge.

The reason nobody caught it for two weeks is the part that actually matters. Other ways of reaching that customer’s site kept working the whole time. Only the short address failed, and every uptime signal I had access to only ever saw the working routes. The account looked completely healthy from every angle I was checking.

What surfaced it was a habit, not a monitor: after shipping a rule like this, run a pass specifically designed to disprove it rather than confirm it. I went looking for exactly the failure mode I’d have argued against if you’d asked me directly, whether any of the strings on that blocklist could also be a real username, and checked the full list against the actual customer table instead of trusting that “obviously a scanner path” meant anything. It took one query to find the collision.

The fix was to remove the clause rather than narrow it. A scanner probe path is a nuisance. A paying customer’s homepage returning an error is not a tradeoff to manage, it’s just wrong, so the obscure signature lost and the customer’s address won. I then checked every other value still on that same rule against the customer table, close to thirty entries, and found no other collisions, which told me this was a one-off coincidence and not a pattern I needed to redesign around.

The part I didn’t expect was finding a second bug while I was in there. A tool I’d been using to summarize what was actually on that rule only understood one kind of match clause and was silently skipping others, which meant the count it reported had been wrong for a while and made the rule look shorter and more reviewable than it actually was. That’s part of why the collision sat there unnoticed: the tooling I trusted to tell me what the rule contained wasn’t telling me the whole rule. Fixed that in the same sitting, since a summary tool that quietly drops entries is worse than no summary tool at all.

The lesson isn’t “don’t block scanner paths.” It’s that an exact-match rule on a short string is only as safe as the space of real values it might also equal, and “I checked it wasn’t getting legitimate 200s” answers a different question than “I checked it wasn’t someone’s actual address.” Those look like the same check. They aren’t, and the gap between them is exactly where a real customer sat, quietly locked out of their own site, for two weeks.