We told a client their site would be live on the new relay by Friday. Friday came and went. Their site was fine, technically, still resolving to the old setup, but we had made a promise about a date we did not control, and it was awkward to walk back.

That was the moment I stopped writing status updates about outcomes and started writing them about state.

The setup

We are mid-migration on a project moving domain handling for one of my contracts, a classic-ASP web app, onto a public relay running on DigitalOcean. The relay runs Caddy with on-demand TLS: it does not pre-provision certificates, it mints one the first time it sees a real TLS handshake for a domain that is on its allowlist. No handshake, no cert. No cert, no site on the new path.

That detail matters more than it sounds like it should, because it means the relay’s job is only half the story. The app can push a domain onto the allowlist all day long. But nothing actually completes until traffic for that apex domain shows up at the relay, and traffic only shows up once the customer’s DNS points there. We do not control that step. The customer’s registrar, their delay in updating a record, their IT person who is out until Monday, that is the real bottleneck, not our infrastructure.

We had been talking about the migration like it was purely an engineering timeline. It was not. It was an engineering timeline gated by a customer action we could not schedule.

Where the honesty broke first

The allowlist itself taught us this the hard way before the messaging did. The sync script that kept the relay’s allowlist in step with the database, Sync-RelayAllowlist.ps1, was written to treat the relay as something that should always exactly match the DB: whatever the DB said, the file should say. That sounds correct. It is not.

The DB knows what domains exist. It does not know which of those domains have actually completed the customer-side DNS cutover. So a sync that replaces the allowlist wholesale can silently drop a domain that is mid-migration, live on the old path, waiting on the customer, but not yet reflected as current in the DB’s view. Drop it from the allowlist and the next handshake attempt fails outright. A domain that was fine yesterday breaks today, for a reason that has nothing to do with anything we changed on purpose.

The fix was to stop treating “matches the DB” as the goal. The allowlist became append-only, a union of what the DB knows plus a separate supplemental file for domains we own the timeline on. Removal became a deliberate, reviewed action, never an automatic side effect of a routine sync. As the design note for the fix put it: removal is the direction that breaks things. Additions are cheap and safe. Deletions are the ones that need a human to mean it.

Renaming the status

Once we’d fixed the mechanism, the messaging problem was still sitting there. We were still telling customers things like “you’ll stay online through the switch” and “expect this done by end of week,” promises that depended entirely on an action outside our system.

So we changed the vocabulary. Instead of a completion date, we started reporting two states: cert ready, or awaiting customer. Cert ready means the relay has the domain on its allowlist and Caddy has successfully minted a certificate, our half is done, full stop. Awaiting customer means exactly what it says: we are done, and the next handshake that finishes the job depends on them updating a DNS record we cannot touch.

That distinction turned out to be more useful than any timeline we’d been giving. It told the customer precisely what to check on their end. It told our own team, at a glance, whether a stalled domain was an infrastructure problem or a customer problem, which is a very different Tuesday depending on the answer. And it stopped us from absorbing blame for delays that were never ours to control.

Why this generalizes

The instinct to promise an outcome is understandable. “Your site will be live by Friday” sounds more reassuring than “the cert is ready whenever you update your DNS.” But the first one is a bet on someone else’s timeline dressed up as a commitment, and when it slips, it reads as a broken promise even though nothing on your end failed.

Reporting state instead of promising outcomes does not feel as confident in the moment. It is more honest about where the actual dependency sits, and it holds up a lot better a week later when the customer finally gets around to their DNS change and the cert is, in fact, already sitting there ready. The work was done on time. We just stopped pretending we owned the last step.