Nearly a thousand of our TLS certificates for customer-facing sites were sitting on a validation method that needed a fresh DNS record from the customer’s own domain every 90 days, and nothing in place would have noticed a specific one about to fail before the renewal actually did.
How it happened
Months earlier, a one-time bulk migration onto a CDN’s managed-certificate product had registered close to a thousand hostnames using the older of two validation methods, the kind that checks a special record sitting in the customer’s own DNS. Every hostname added since had gone in on the newer method by default, the kind the CDN answers itself at its own edge, with no DNS record required from anyone. Nobody had chosen the older method on purpose for that batch. It was just what the migration tooling happened to write, and it sat there working quietly for months, because a certificate on that method only actually needs the DNS record present at the moment it renews.
Why the existing safety net wouldn’t have caught it
We already had a script that watches for certificates stuck mid-renewal and repairs them automatically. It only acts on a hostname it can already see failing, one sitting in a stuck or timed-out state. A hostname on the older method that hasn’t hit its 90-day mark yet doesn’t look stuck to anything. It looks exactly like every healthy certificate, right up until the DNS record it needs isn’t there anymore and the renewal actually fails.
What made it visible
The CDN itself surfaced it, sending “action required” notices as each 90-day window approached. Reading a handful of those closely turned up the real shape of it: a batch, not a one-off, all sharing the same migration origin, all coming due around the same season.
The fix, and the number that mattered
Every hostname on the old method got flipped to the new one in a single pass. Zero of the flips failed. The estate afterward was clean, all of it on the method that needs nothing from any customer, ever, at renewal.
The gap that made this possible in the first place
The reactive script and this fix solve different problems, and neither replaces the other. One repairs a renewal that has already broken. Nothing before this existed to notice a batch quietly approaching that break before any of them actually hit it. That’s what got built alongside the fix: a watchdog that reads the whole estate on its own schedule and says something the moment a certificate is close to expiring without a completed renewal, a hostname has drifted back onto the old method, or an onboarding never got a certificate at all, instead of waiting for a vendor’s warning email or a customer to notice first.
The lesson underneath it
A fix that repairs failures after they happen and a check that notices trouble before it happens are not substitutes for each other, and a system can look healthy for a long time running on only the first one. The certificates that had renewed fine every 90 days for months looked exactly like the ones that never would have, right up until the day one of them didn’t.
AI Skills
Use this lesson with the AI assistant you already use
A script that repairs failures after they happen only ever sees the ones that already broke. It can't say anything about the ones quietly approaching that same wall.
Paste the prompt, share only the context needed to answer it, and treat the result as a draft for your review. Do not include confidential information or let an AI assistant make changes without your approval.
Optional: for a visual report and saved memory, run /dxdev first.
Don’t have it? Get it at dxdev.com/skills/dxdev. The prompt works without it.
dxdev LESSON · paste into your AI coding agent
LESSON: A Reactive Fix and a Preventive Check Are Not Substitutes for Each Other
SOURCE: dxdev.com/blog/2026-08-06_the-batch-aging-out-with-no-alarm
WHAT HAPPENED: A one-time bulk migration months earlier had registered close to a thousand certificates on a validation method that needs a fresh DNS record from the customer's own domain at every renewal, instead of the newer method that needs nothing from anyone. Nobody chose that on purpose; it was just what the migration tooling wrote at the time, and it worked quietly for months. An existing script already repaired certificates that got stuck mid-renewal, but it only acted on a hostname it could already see failing. A certificate sitting fine on the older method, months from its next renewal, looked identical to a healthy one right up until the DNS record it needed wasn't there. The batch was found via the vendor's own expiry-warning emails, flipped to the newer method in a single pass with zero failures, and a separate watchdog was built to notice this class of drift (an approaching expiry with no completed renewal, a hostname drifting back to the old method, an onboarding that never got a certificate) before any vendor email or customer ever had to say so.
THE RULE: A script that fixes things once they're already broken and a check that notices things before they break are answering different questions, and having one does not mean you have the other. If your only signal for a slow-building class of failure is "did this specific thing already fail," you will not see the batch quietly approaching that same failure until members of it start hitting it one at a time.
CHECK MY CODE, then report PASS or FAIL with file:line for each:
1. Any auto-heal, retry, or repair job that only triggers off an already-failed or already-stuck state, with nothing checking for things approaching that state.
2. Any one-time migration or bulk-provisioning step whose output was never re-verified later against the setting it was supposed to produce.
3. Any recurring dependency (a cert renewal, a token refresh, a scheduled resync) where the interval is fixed but nothing tracks how close each individual instance is to its own deadline.
THEN PRINT: a table (check, PASS/FAIL, evidence, fix) + a verdict (applies / partially / OUT_OF_SCOPE / no) + the single most important next action.