The status alert said the relay had been down for 22 minutes, and the relay was a box I thought I understood.
It terminates TLS for the bare domain of every customer site. That is 1,295 domains, listed in a flat file at /opt/relay/apex-allowlist.txt, all served by Caddy on a 458 MB droplet. The www hosts run through a separate path and were fine. Only the bare apexes were dead. Port 22 answered and port 80 refused the connection, so the box was up and Caddy was gone. The kernel OOM killer had taken it, and nothing brought it back.
Twenty-two minutes to a restart
I asked the agent working the incident for the first move, and it wanted a go-ahead before touching prod. It also offered to open a ticket for the root cause. I said “no tix, fix now.”
systemctl start caddy, then a check that 80 and 443 were listening, then curls against real customer apexes. They returned 301 to www with a clean TLS verify. Outage total: 02:01 to 02:23 UTC.
The root cause was two missing things on a small box:
- No swap. The box had none, so any memory spike went straight to the OOM killer.
- No restart policy. The packaged Caddy unit has no
Restart=, so a killed process stays killed.
The fix was a systemd drop-in at /etc/systemd/system/caddy.service.d/10-restart.conf:
[Unit]StartLimitIntervalSec=600StartLimitBurst=20
[Service]Restart=on-failureRestartSec=5sI set the start limit generously on purpose. A crash loop on this box is still better than a permanent outage. I added a 1 GB swapfile with vm.swappiness=10, so swap works as a cushion and not a paging target. I applied the same hardening to the second droplet and resized the relay for real headroom. Swap on a 458 MB box only delays the next kill.
Killing it on purpose
A config file is a claim, so I killed the thing. I recorded the main PID, sent SIGKILL, and polled port 443 from my machine for the gap. Refused, then back after 5.9 seconds. The PID had changed and NRestarts read 1.
The kill command printed Failed to send signal SIGKILL to auxiliary processes: Invalid argument. That looks like a failed test, but the main process died anyway and the restart policy fired. The polling loop gave me the real answer. I trust the port over the exit message.
The line in my incident note
In my own incident note I wrote: “The alert did its job, probe-backed down, paged on the first build, exactly as designed.”
That was true of the alert, and I let it stand in for the whole system. A green verdict on the monitor told me nothing about the jobs the monitor doesn’t watch. The relay’s allowlist is refreshed by a sync job, and while I was reading the allowlist file to size the incident, I found that job had not run since June. That is about two months of silence. No alert fired, no error was logged, and nothing was down. The file still had 1,295 lines, so it looked fine.
A silent failure and a loud one are different problems. A crash that leaves a service dead is loud. Fix it with a restart policy and you are done. A job that stops running and leaves stale-but-plausible output behind needs a check that asks whether the output is fresh. The restart policy I had just proven covers the first kind and does nothing for the second.
A ticket filed where the job never lived
I had said no ticket for the outage, but this one deserved a follow-up, so I filed one for the allowlist sync in the platform’s tracker.
Later that link was gone. The ticket had never belonged there. The sync never lived in the platform repo. It runs as a scheduled task on one of my own machines, outside anything we deploy. I had filed the work in the place I assumed it lived, and the assumption cost me the trail. Whoever followed that ticket would have searched a repo that has never contained the job.
So the follow-up went stale for the same reason the job did. Both lived outside the system that would have noticed. The fix for the ticket was to strip the dead link and record the real location. The fix for the job is what I still owe: a freshness check that fails loudly when the allowlist’s last successful sync is older than a day.
Rebuilding the alert path
The outage also exposed how alerts reach me. The old path was one channel, and I was the single point of delivery. I rebuilt it into a ladder that says out loud which rung is speaking:
- My own agent for that venture, checked through a status beacon that says whether it is listening.
- That venture’s always-on bot, announced as the backup.
- SMS, announced as the last resort.
All three rungs were exercised end to end, including a real text. Status pages now send from the agent when it is listening and fall back with the reason stated when it can’t act.
The ladder guards against my not hearing an alert. It does nothing about an alert that never fires. The allowlist job proved that after the fact. The restart policy covers the crash you can see, and the freshness check has to cover the run that quietly never happened. Six seconds of downtime became the cost of a crash. Two months of stale data was the cost of a job nobody was watching.