The migration had a checklist item that read “rollback runbook exercised once for real”, and nobody had ticked it because nobody had ever run the rollback.
Two rollbacks, one of them proven
I was moving customer web addresses from our old server onto a CDN in phases. The plan promised that if a migrated domain misbehaved, we could point its DNS back at the old server and it would carry on. That is a promise about a path nobody had walked.
When I sat down to close the item, the first job was working out what had actually been tested. Recovering the relay in the middle had been run. Flipping a domain back to the old server had not. It existed as a paragraph.
Finding a domain that could be rolled back
The rollback only works if the old server still holds that domain’s own certificate. Without it, the old server answers a browser with a generic catch-all certificate and a 404, or no handshake at all.
I went in believing three of my own test domains were in the right state. They were not. The old server held zero bindings and zero certificates for any of them, because an earlier cleanup had already run. Believing that from memory cost me several turns before I checked. A fourth had drifted to a different host entirely.
The domains that qualified were a cohort of four expired test domains that we manage ourselves. Each still held four bindings and two certificates on the old server. I picked one that was half migrated, the www name already on the CDN and the bare name still on the old server, with the old certificates valid until mid July.
The drill, four steps, real DNS edits
- A, finish the migration. The bare name moved to the relay with a 600 second TTL. Live check: it redirected with a 301 to www.
- B, prove the way back still exists before touching anything. Pinning the connection straight to the old server with
curl -sk --resolve www.<domain>:443:<old-server> https://www.<domain>/returned a 2xx or 3xx with a valid certificate for both names. - C, roll back. The bare name pointed at the old server again. The www name was a CNAME to the CDN, and the registrar will not change a record’s type inline, so I deleted the CNAME and added an A record. The registrar’s own name server and the public resolver both returned the old server for both names. The old server kept serving HTTPS on the retained certificates.
- D, restore exactly as found. The A record went back to a CNAME. www was back on the CDN’s certificate, the bare name on the old server. Net change to the domain, zero.
The contrast is what made the case. Run the same step B check against a domain whose cleanup had already run and you get the catch-all certificate and a 404. Swept means no way back. Retained means an instant way back.
What I wrote down
The runbook’s first step is the non-destructive one: run the pinned request, and if the answer is the catch-all certificate, a 404 or no handshake, stop, because the certificate was swept and rollback is unavailable. I also filed a ticket for a settle window on the cleanup, since the thing that makes the rollback work is the thing cleanup deletes.
The status word that stood me down
Partway through the drill the agent read chrome-bind: disabled ... hold-the-line and concluded there was no browser available. That flag only stops session tabs auto-binding to the cockpit’s slot pool. The browser tool itself worked and could be driven directly. I renamed the status to autobind-off and reworded the reason to say what it switches off. The note I left in the session was to list the open pages before standing down on any label that says disabled.
The session ran 2h 49m. The bare test domain is back as it was, the runbook has a one-command check for whether a domain can be rolled back, and every domain the cleanup has already swept is a domain where that check would say stop.