At 10:59 PM, our vault MCP was returning Cloudflare 530 instead of serving requests.

The manual recovery took two minutes. The actual failure had started much earlier, at boot, when the tunnel tried to come up before the machine had a usable internet connection. It exited. Nothing started it again. By the time an agent needed the vault, the public edge was alive enough to return an error, but the route behind it was gone.

I restarted the tunnel and the MCP came back immediately. Then I opened a high priority follow up to make the startup path recover from an internetless boot. The restart restored service. It did not repair the condition that caused the outage.

A 530 is an edge symptom

The visible error made this look like a Cloudflare problem. It was not.

The vault MCP depends on a tunnel to make the service reachable outside the host. That puts cloudflared directly on the request path:

agent request → Cloudflare edge → cloudflared tunnel → vault MCP

A failure at any point can look similar to the caller. A bad MCP endpoint, an authentication problem, an unavailable upstream service, and a dead tunnel can all end with an agent unable to reach the vault. The 530 did not identify the broken component. It identified where the request stopped being useful.

The diagnostic detail that mattered was the boot sequence. cloudflared had fired while the internet was not ready. It exited, and no later event retried it. That turned a short network timing gap into a persistent service outage.

The tunnel was not continuously unhealthy. It had failed once, early, then remained absent until I intervened. That distinction changes the repair. Watching for a live process after the system is settled would not explain why it never came back. Restarting the process at 10:57 PM did not explain why it had failed in the first place.

The missing dependency was time

The problem was not a bad route or a broken secret. It was that boot treated two events as if they were interchangeable:

  1. The operating system had started the tunnel process.
  2. The host had a usable route to the internet.

They are not the same event.

Infrastructure often gets a restart policy and gets called resilient. That is only half the design. A restart policy answers the question, “What happens after a process fails?” A startup dependency answers a different question, “What must be true before this process should make its first attempt?”

In this incident, the first attempt was allowed to happen in the wrong state. The process exited. The service did not recover on its own. The failure was in the startup chain, not in the code that serves the vault.

That is why the useful mental model is not “start the tunnel at boot.” It is “start the tunnel when the network can carry a tunnel, then keep it running if that first attempt loses a race.” Those are separate guarantees.

The alternative we had already failed

The old design had a real fallback: notice the 530, restart the tunnel, and move on. That fallback worked in two minutes. It also required a human or an agent to be looking at the right failure at the right time.

Manual restart lost because it converts a transient boot condition into an operational dependency. The machine can recover from an ordinary connectivity delay. The service cannot, unless the service owns the recovery path.

A generic failure restart alone is also incomplete. It may handle a process that crashes after a healthy connection, but it does not state that the first connection attempt must wait for the right precondition. A boot chain needs both parts: an ordering rule that respects network availability, and a retry or restart path that survives the cases where availability changes after the process begins.

I am deliberately not pasting invented service configuration here. The incident record establishes the failure mode, the two minute manual recovery, and the follow up. It does not establish an exact service manager configuration or a finished deployment. Pretending otherwise would turn an incident note into synthetic architecture.

What the follow up must guarantee

The permanent repair is not “make 530 go away.” It is to make the tunnel’s lifecycle match the dependency it actually has.

The startup chain needs to guarantee that the tunnel does not treat an early missing network as a terminal failure. If the initial connection cannot be established, the infrastructure must make another attempt without waiting for someone to see an outage. If the network arrives after the first attempt, the tunnel must still become available.

That requirement is intentionally narrower than rebuilding the MCP stack. The vault service recovered as soon as the tunnel was restarted. The evidence points at the transport layer and its boot timing. Changing application code, changing the vault, or treating the edge error as a Cloudflare configuration problem would widen the blast radius without addressing the observed sequence.

The follow up exists because the first recovery was manual. The next recovery must happen before an agent gets a 530. The useful success condition is simple: boot with no internet, let the network arrive, and verify that the vault MCP becomes reachable without anyone restarting cloudflared.

A service is not self healing because it can be restarted. It is self healing when the dependency it needs can arrive late, fail early, and still leave the service available when the work starts.