55 certificates were stuck, and every alert line read ssl status 'pending_validation' with a live cert. I spent the first stretch of that morning treating it as the thing the watchdog was built to catch. Eleven customer sites had no HTTPS at all, and they sat inside that same list looking exactly like the other 44.

What the watchdog believed

The watchdog exists because Cloudflare for SaaS issues 90-day DV certs for every customer hostname and renews inside a roughly 30-day window. Renewal needs domain control validation. Our June bulk migration registered 956 hostnames on txt validation, which needs an _acme-challenge record in the customer’s zone. The registrar’s DNS API appends instead of overwriting, so stale tokens pile up and renewals fail silently. On August 6 we moved all 956 to http validation, where Cloudflare answers the challenge at its own edge. The watchdog was the guardrail behind that change. It flagged four kinds of finding: expiring, stuck, txt-drift and stalled.

Every one of those kinds reads certificate state from the Cloudflare API. None of them asks where the hostname’s DNS points.

That is the hole. A certificate only matters if Cloudflare is the thing answering for the hostname. If DNS points somewhere else, the SaaS record is inert. Its cert is unused and HTTP validation can never complete, because the challenge request never reaches the edge. So pending_validation on a hostname mid-cutover is normal migration progress. The same status on a hostname whose DNS still points at us is also what you see when the origin cert behind it has been deleted.

Reading the alert the way the tool’s header told me to

I took the alert as a real renewal problem. That was the natural reading, since the tool’s own header says stuck means a renewal is in flight and failing. I went at the pile as 55 renewals to unstick.

It cost me the first hour of the morning on the wrong question. Renewal state was a symptom for all 55, and for 11 of them a misleading one. The migration cleanup had retired the origin certificates for those sites while their DNS still pointed at us. Cloudflare had nothing to validate against, and our origin had nothing to serve. Visitors got a TLS failure.

What gave it away was checking the sites the way a visitor would. For the 44 that were mid-cutover, DNS resolved somewhere else and the site loaded fine. For the 11, DNS resolved to us and the handshake failed. The certificate status was identical across both groups, so it could never have told them apart.

I restored all eleven origin certs and verified each one over HTTPS.

Teaching it to resolve first

The fix goes in a function called apply_dns_reality. It runs after the normal findings are built and reclassifies only the certificate-state kinds. txt-drift is a configuration fact that holds whatever DNS says, and stalled already means no cert exists, so those pass through untouched.

For each remaining finding it does the following:

  1. Resolve the hostname, and compare the answer to the addresses EDGE_CNAME resolves to right now. That live lookup is the primary test.
  2. As a backstop, check the answer against a snapshot of Cloudflare’s published IPv4 ranges (CF_RANGES). Without it, a rotation of the zone’s addresses between two lookups would flag the entire estate as off-edge at once.
  3. If it resolves to the edge, the cert finding stands and keeps its original kind.
  4. If it resolves elsewhere, open a real TLS connection to wherever DNS points and see whether HTTPS works. On success, the finding is downgraded to off-edge, an informational kind that never pages. On failure, it becomes origin-https-down.

The lookups run in a thread pool of 8 workers, since 955-odd handshakes serially is a slow way to run a cron job.

Ranking matters too. The report used to sort by a local dict inside the function. I hoisted it to a module constant:

RANK = {"origin-https-down": 0, "expiring": 1, "txt-drift": 2, "stuck": 3,
"off-edge": 4, "stalled": 5}

origin-https-down is the only kind that means a customer cannot reach their site over HTTPS right now, so it sorts above expiring. Before this change, that condition had no kind of its own and no rank at all.

After the change, the 55 lines split the way they should have from the start. Forty-four read as migration progress that will settle. The other 11 read as an outage, at the top of the page, with a name that says so. I also filed a ticket for the guard hole that let cleanup retire origin certs for hostnames still routed to us, because the watchdog fix only makes the failure visible. It doesn’t stop the deletion.

A count of certificates is not a count of broken sites

When a monitor reports 55 of something, I now ask what population it is counting and what it never looks at. This watchdog counted certificates. What customers experience is a route ending in a handshake. Those overlap most of the time, which is why the tool looked correct for four days, and they diverge exactly when a migration is half done.

Since the tool now says which sites are mid-cutover and which are dead, a weekly run can tell us the difference on its own.