The alert never fired. What fired instead was a teammate asking twice, five days apart, whether his preview build had actually gone out.

It had not. Every push to that branch was hitting our CI, running clean, and reporting success. The plan itself had been switched to disabled sometime in the prior week. Our CI system does not reject a build against a disabled plan. It accepts the trigger, marks the job as skipped, and returns green. From the pusher’s side there is no difference between “deployed” and “silently discarded.” Same webhook response, same commit status check, same nothing in the CI tab that would make you look twice.

I found it by working backward from the symptom, not the log. The teammate said the preview site looked stale. I diffed the commit hash the site was serving against git log on the branch and they were five days apart, right at the point his last real deploy had landed. Then I went looking for the failure and there wasn’t one. No failed run, no error email, no red X anywhere in the build history, because there was no build history for those five days at all. The plan just stopped showing up in the list of things that ran.

That’s the part worth sitting with: I checked the CI system’s audit log for who or what disabled the plan and there was nothing to check. No entry, no timestamp, no actor. Either it predates what the instance retains, or the disable happened through a path that doesn’t log to that table at all. I re-enabled the plan, kicked a manual build, confirmed the site was now serving his actual latest commit, and closed that part out. The cause is just gone. Not redacted, not buried in a log I didn’t look at hard enough. Unrecoverable.

My first move, before I went hunting in the CI system’s UI, was to assume it was a caching problem on the preview host itself. Preview environments in this stack sit behind a CDN layer that we’ve had misbehave before, serving a stale build after a real deploy because of a cache key that didn’t bust cleanly. So I bounced the CDN cache for that host, hit the URL again, watched the same five-day-old commit hash come back, and lost maybe forty minutes chasing a cache invalidation that was never going to fix anything, because the new build had never been produced to begin with. The deploy artifact for the last five days simply did not exist on the target. Cache had nothing to invalidate.

That wrong turn is what pushed me to stop trusting “the pipeline says success” and start comparing what the pipeline reported against what was actually on disk on the target host. Once I did that the gap was obvious in about two minutes: five days of green checkmarks with zero corresponding artifacts.

The fix for the immediate problem was one click: re-enable the plan. The fix for the actual problem is a ticket we filed to track the monitoring gap, because the real failure isn’t that a CI plan got disabled. Config drifts, people fat-finger toggles, that’s going to keep happening. The failure is that our monitoring watches for jobs failing and has no opinion on a job that stops running entirely. A disabled plan doesn’t fail. It just stops existing as a thing that could fail, and every downstream check that says “no failures reported” reads that as healthy.

What we filed isn’t “alert when the build fails,” we already have that. It’s a heartbeat check: has plan X produced an artifact with a timestamp newer than N hours, independent of what the CI system’s own status API claims. If the last successful artifact is stale, page someone, regardless of whether the plan even ran. That’s the only version of this check that would have caught it, because from the CI system’s own point of view, nothing had gone wrong. Nothing had gone at all.