At 8:36 on July 20, I went in to answer whether five client-managed domains would break at an August 7 cutover. Two would. Three would not. Before I closed the question, I found 55 Cloudflare custom-hostname registrations that no longer existed.
They had not failed visibly. They were gone.
That distinction matters. A failed registration gives you something to inspect: an error state, a retry, an API response, a bad certificate. A silently deleted registration gives you an empty space. Our dashboard could show the registrations it still knew about, but it could not show the ones Cloudflare had removed when an endpoint changed. Months can pass in that gap, right up until a customer hits the broken path before we do.
I fixed the 55. The more important work was making the same class of disappearance observable every day.
The inventory was lying by omission
The original question was about a small cutover cohort. Five domains had been partially migrated, so I traced the origin path for each one and determined which setup would be live on August 7. That produced the immediate answer: two were at risk, three were safe.
But that investigation forced a comparison we had not been making continuously. We have a product-side set of domains that should have a Cloudflare custom hostname. Cloudflare has its own set of registrations that actually exist. Those are not the same thing as a status field inside our application.
For this configuration, the expected Cloudflare object is www.<domain>. The apex is not a Cloudflare custom hostname. That detail is small enough to get lost in a migration script and large enough to turn a correct audit into a stream of false positives.
The invariant is equally small:
expected Cloudflare hostnames - actual Cloudflare hostnames = registrations to investigateThe result was 55.
Nothing in the dashboard was wrong in the ordinary sense. It was reading existing registrations. The error was in treating that view as an inventory. A dashboard is often a projection of the records that survived. It is not proof that the remote system still contains every record we expect.
That is the failure mode with cloud APIs that delete state behind your back. The user interface has no row labeled “this used to exist.” The absence has to be computed against an independent expected set.
I repaired the write path first
The immediate repair was a bulk re-registration of the 55 missing Cloudflare entries. I verified it against the live system. That restored the missing objects, but a bulk repair alone would have been an expensive way to manufacture the next incident.
The durable fix went into CFWrite, the path that manages the registration. It now handles a deleted registration as a purge and recreate operation. I did not try to massage a stale local record back into a healthy state. When Cloudflare has removed the remote registration, the safe action is to discard the stale registration state and build the remote object again.
That choice was deliberate. We could have kept the old status and attempted an in-place update. It would have preserved a record for the dashboard, but it would not have restored the object that Cloudflare had deleted. We could also have patched the 55 and declared the incident closed. That would have made the data look clean for one day while leaving the detection gap intact.
I also considered making the staff page recreate missing registrations when somebody opened it. That lost for two reasons. It would turn a read into a side effect, and it would make staff activity, rather than the remote system’s condition, determine when a broken registration was repaired. A customer should not need a staff member to visit a screen before their domain becomes healthy again.
The staff work still mattered. The dialog and help content now distinguish customer-fronted domains. That classification belongs before the comparison. If a customer owns the front of a domain, it must not be added blindly to the expected Cloudflare set. Otherwise the guard would be technically active and operationally useless, because it would train staff to ignore its false alerts.
A guard is not a cron line on a workstation
The repair path handles a known missing registration. It does not prove that tomorrow’s endpoint change will leave the correct remote state behind. For that, I added CF-RegistrationGuard, a daily sweep that runs the expected-versus-actual comparison and catches the difference before the dashboard has a chance to normalize it away.
The first scheduled task landed on my workstation. It ran, which proved the task could execute, but it was the wrong deployment. A workstation is not infrastructure. It sleeps, reboots, loses a session, and belongs to a person rather than the system. I removed that version and put the daily guard on the origin box, where the rest of the domain automation lives. I verified the corrected deployment by count.
That placement is part of the design, not deployment housekeeping. A continuous check only counts as continuous when its host, schedule, and execution context are part of the system it is checking.
The daily job also beat an event-only design. It would be tempting to validate only when CFWrite changes an endpoint. That is useful, but it cannot detect drift that occurred before the new code shipped, and it makes the same remote call responsible for both changing state and judging whether the provider retained it. The scheduled sweep approaches the problem from the other side. It asks, independently, what we expect today and what Cloudflare actually has today.
The verifier needs a verifier
I did not leave the guard as an untracked script on the server. CF-RegistrationGuard and the companion detector are part of the engine file list and the committed deployment manifest. A later parity check confirmed that the engine files on the origin box and in the repository were byte-identical.
That is a separate but related lesson. A daily verification script is only useful if the production machine is running the version you think it is. Otherwise you can build an excellent guard, commit it, and still operate a server with yesterday’s assumptions.
The incident started as five awkward migration cases and ended with 55 repairs, a corrected write path, a daily reconciliation job, customer-fronted domain handling, and a deployment trail that can be checked again later. None of those pieces is sophisticated on its own. The important move was refusing to let the provider’s dashboard-shaped absence define reality.
Cloud services are good at reporting the state they still hold. The systems we build around them need to be equally good at reporting the state that disappeared.