The publish log for a 23 hour 53 minute run showed zero errors, and the live post count went from 554 to 731. That is 177 posts, and I still opened every one of them on the live site before I trusted the number.
What a clean log measures
The pipeline drafts a post, rewrites it, revises it, and pushes it. Each step exits 0 or it doesn’t. When every exit code is 0, the log says the job ran. It says nothing about whether the page a reader sees matches what the job thought it shipped.
So the spot check was a plain loop. Take a slug from the run, load it on the public site, and compare it against what the log claimed. Two follow-up tickets came out of that, one against state tracking and one against authorization. I can show evidence for only one of them. The state tracking evidence is sitting in the ordinary daily logs. For the authorization one I don’t have a clean mechanism yet.
State tracking was already leaking in the daily logs
The nightly backfill job ran at 2:02 AM the same day and reported that nothing was published or held. Buried in its summary: six candidates turned out to be duplicates of posts that were already published, so it deleted them. A job that thinks it has six new posts to write, and discovers after the fact that they exist, has no reliable record of what it has already done. It recovered that time. Under a 177-post burst, that lookup gets many more chances to be wrong.
The nightly sweep tells the same story from another side. Its resync line that night read:
resync: 3 active ticket session(s), 3913 zombie thought folder(s), 4 stuck-active manifest(s)needs-close: 219 still need a human /close3,913 folders and 4 manifests all claiming to be live while nothing was running is what state drift looks like when it’s counted. None of it threw an error. Each record was internally consistent and only wrong against reality.
The same sweep’s checkpoint step logged this:
checkpoint: committed 0 repo(s): HELD (live session working there): vault 262 file(s)That hold exists because of a bug I fixed the same day. The sweep used to commit other sessions’ in-progress work. Now it refuses to commit a repo while a live session is working in it. The blog session was the live one, holding 262 uncommitted files. It’s a small rule about ownership of state, and it only exists because a log full of successful commits hid the fact that they were the wrong commits.
Why the run was on a different account
The night before, the blog jobs were blocked because the account they ran under was nearly out of usage. I moved the work onto the main account through Sunday and let it burn through the backlog in one long session. That kept the run alive, but it also meant one run touched far more posts than the normal daily job, on a different account than the one the pipeline was built around. That is where authorization assumptions are most likely to break. I filed the ticket and did not pretend to have it solved.
Closing the gap between “exit 0” and “page is right”
I did not want spot checks to stay a manual ritual, so the same day I changed the site crawl. It now flags any page that shows error text or is missing its main nav, and it only alerts on new failures, so a known-bad page doesn’t page me forever. A new 10-minute task runs a 13-page production check after every push, and the weekly crawl is back on.
Here is the state of that fix. A real push triggering the 13-page check has not been observed yet. A real chat alert has only been mock-tested. Its first weekly run will also list four crawl-artifact pages as new, and those will be noise. So the tool that exists to stop me trusting logs is itself only verified by a mock, and I’ll trust it after I watch it fire on a real push.
Until then, the rule I run by is simple. A log tells me the job finished, and only the live page tells me what got published. For 177 posts, that meant 177 loads, and two of the follow-ups I filed came from those loads.