The emails started as “[Action required] Validate
956 of 1,265
Every customer domain we host rides a CDN as a custom hostname with a 90-day certificate. Renewal needs the domain to be validated again, and there are two ways to do it. With the HTTP method the CDN answers the challenge at its own edge, so nobody has to add anything anywhere. With the TXT method the customer’s own DNS needs a record at every renewal.
Normal onboarding had always used HTTP. The June bulk migration had registered 956 hostnames on TXT, and nobody had chosen that for them. Everything registered through normal onboarding was already on HTTP. The 956 certificates all came due in September, which is why the notices started arriving in August. Each one needed a record in the customer’s own DNS, and a renewal with that record missing fails silently.
I moved all 956 to HTTP. The estate now reads 1,265 of 1,265 on HTTP and no customer was contacted. The work log closes that stretch, watchdog included, at 10:49 PM, 4h 30m after it started.
The watchdog
Fixing the batch does not stop the next email, so I built a certificate watchdog with three commands:
checksettles one “action required” email. It reads the CDN’s record and the certificate the edge is actually serving, then says whether the email needs anything. Most of those emails are stale by the time they arrive.scanflags a certificate inside 14 days of expiry that has not renewed, a hostname drifting back to TXT, a renewal stuck in flight, and an onboarding that never got a certificate. With--notifyit DMs me on Discord with an SMS fallback, deduped by a signature so a standing finding does not page every day.inventorydumps the whole estate to JSON and CSV.
It runs daily at 07:20 as a scheduled task and costs nothing in model spend, since it is the CDN’s API plus the notification path.
One false page, then 55
Two days later it paged me that one hostname was stuck. Its renewal had already completed. check said healthy and scan had never asked. So scan --notify now re-verifies every finding against a fresh fetch just before paging and drops anything that has healed.
The bigger miss came four days after launch. One morning scan paged me 55 hostnames as “stuck”. None were stuck. All 55 were domains still waiting to be cut over to the CDN, with DNS still pointing at the old server. Their CDN certificates sat unused and could never finish validating, because HTTP validation needs the traffic to arrive at the edge. That is what pending looks like mid-migration, and nothing a visitor saw was wrong.
Inside that batch was the thing I would actually have wanted a page for. Eleven domains had lost their bindings and certificates on the old server in the migration cleanup, while their DNS still pointed at it. Loading their https address returned a connection reset while the plain http address still served. Certificate state on the CDN could not tell those eleven from the rest of the batch. Where the hostname resolves could.
scan and check now resolve the name first and sort each finding:
- off-edge: not on the CDN yet and HTTPS is fine wherever DNS points. Reported every scan, never paged, and the muted count is printed and logged so a quiet run is never mistaken for a clean one.
- origin-https-down: not on the CDN and the host cannot complete a TLS handshake. This outranks everything, because https is dead for that customer right now.
The eleven had their bindings and certificates restored. The recheck before paging runs the same classification, otherwise a reclassified finding could never match its earlier self and every one would look healed.
I began the day thinking the risk was 956 certificates. The bigger risk turned out to be that my alarm could not tell an unused certificate from a dead site.