At 7:10 a.m., the apex relay health monitor fired. It fired again at 7:30. By 3:38 p.m., we had moved 25 brand domains, including the primary domain, off the old origin server and onto the apex relay and Cloudflare. The sites stayed up. Mail stayed intact.

That was not a normal DNS change. We have a long running partnership business with staff and customer facing systems, but we do not have a dedicated operations team waiting to absorb an infrastructure migration. The same people responsible for the product and customers have to make the infrastructure change legible enough to operate.

The difficult part was not adding 25 domains to Cloudflare. It was deciding which system would notice each change first, and making every transition reversible before the next one began, while the primary domain and mail for a live customer-facing business stayed exposed to every single step of it for over six hours.

The alerts were a boundary, not a diagnosis

The morning began with timeout alerts from the apex relay. At the same time, a customer reported that a domain was not working, and the www CNAME was not pointing where it should. Those facts looked connected. They were not evidence of the same failure.

A relay health alert tells me that a monitor cannot complete its check. A bad www CNAME tells me that one hostname is not following the intended routing path. Neither result tells me, by itself, whether the problem is the old origin, Cloudflare, the relay, the DNS record, or the monitor’s own target.

That distinction changed the work. I did not start by trying to silence the alerts or by making a broad nameserver cutover. I started by mapping the dependencies that could report success or failure: the apex hostname, www, the relay, Cloudflare, mail, and the health monitor. For every domain, I needed to know what received traffic, what resolved it, and what observed it.

The output was an inventory of all 25 brand domains and a migration recipe committed to the repository. The inventory was not paperwork after the move. It was the control surface for the move. Without it, a DNS record that looked safely changed could still leave a www alias behind, or a monitor could continue testing a path that no longer represented production.

I refused the one shot cutover

The tempting alternative was a bulk migration: make all 25 DNS changes, turn on the new Cloudflare configuration, then clean up whatever alerts appeared. That is fast only if nothing is wrong. Once it is wrong, the batch has removed the evidence. You no longer know which domain introduced the failure, which endpoint is still serving the old origin, or whether a rollback is restoring the correct state.

The second bad option was to leave the existing health monitor alone. It was still checking personal and expired test domains, and it had been checking those instead of the company domains customers actually use for some length of time I did not have a number for, entirely unrelated to this migration. That is not a monitoring nicety I’m noting in passing. It means a real production outage on any of the 25 company domains, at any point before this migration, could have gone undetected because the thing meant to watch for it was watching something else. That created a false sense of coverage. A green result there would not establish that the company domains customers use were healthy. A red result could send us investigating an obsolete target while the production relay was fine.

I also could have treated mail as outside the migration because the immediate work was web routing. That would have been another bad boundary. The requirement was not just that pages loaded through the new path. The requirement was that moving DNS and traffic did not disturb mail.

So the sequence became the architecture. First, we established the inventory and the known routing path. Then we moved each domain from the old origin to the apex relay and Cloudflare. At each checkpoint, the next system had to be capable of noticing the change: DNS had to resolve to the intended path, the relay had to receive traffic, Cloudflare had to sit in the new route, and mail had to remain intact. Only after the production path was established did we change the health monitor to company domains.

That last step matters more than it sounds. Monitoring is not an independent dashboard that you bolt on afterward. It is part of the routing design. If it checks the wrong hostname, it tells you nothing useful. If it changes before the route is verified, it can add noise while you are trying to understand the deployment.

The rollback was the next checkpoint

The useful definition of zero downtime here was not that no alert occurred. Two alerts occurred before the migration was complete. It was that every change had a known prior state and a bounded next state.

For a domain, the rollback was not an abstract promise. It meant the old origin path still existed until the relay and Cloudflare path had been verified. For monitoring, it meant we did not replace a stale signal with a different opaque signal. We first moved the service path, then pointed the monitor at domains that represented the service customers actually use.

This is why the www CNAME mattered. A domain is not one thing. It is a set of hostnames, records, targets, and observers. If the apex has moved but www is still aimed elsewhere, the migration is only half done. If the health check remains on an expired test domain, the operational view is still half done.

By the end of the 6 hour 24 minute work session, all 25 domains were on the apex relay and Cloudflare, the primary domain had moved without downtime, mail was intact, and the health monitor had been moved off the personal and expired test domains onto company domains. We closed the migration item with the full inventory and recipe in the repo because this should not become tribal knowledge.

The work did not prove that a small team can avoid operational complexity. It proved the opposite. The complexity was already there in the relationship between DNS, routing, mail, and monitoring. Writing down the sequence forced us to make that relationship inspectable.

The next migration will not be safer because someone remembers what happened on June 14. It will be safer because the system now has a record of which component needs to notice each change, what counts as verification, and what can be restored before the next change proceeds.