The history pipeline is the thing that turns my work log into a searchable system. It picks up transcripts, enriches them with thread metadata, stamps each session with an area, builds a calendar index, and writes the result back into the vault so later sessions can reason about earlier ones. It runs every night at 2 AM as a Windows scheduled task that calls a PowerShell script that calls a chain of Python steps.

I built a health dashboard partly so this pipeline could not rot in silence. The command prints a one-screen invariants dashboard: queue depth, enrichment lag, provider mix, beelink reachability, and the freshness of the nightly cron. The point of having it was so I could glance at it and see when something had stopped working before it stopped mattering.

The pipeline had been quietly failing for four days before I noticed.

the morning audit

The trigger was small. I noticed that the calendar index for a recent day looked thin in the dashboard and went to check why. I expected a one-line answer. I ran the pipeline manually with verbose output and watched it walk the steps, and one of them, calendar_day, raised an exception, got caught at the outer driver, logged a one-line summary into .scratch/cron-refresh.log, and the driver continued. The script finished. Exit code zero.

Every other step had run. The calendar step had not. The cron had been arriving at exit code zero every night for four nights with one step quietly returning nothing.

The dashboard read the script as healthy because the script said it was healthy. The script said it was healthy because the driver had decided, somewhere a year ago when this code was first written, that one step failing should not bring down the whole pipeline. That was reasonable in the moment. The driver caught the exception, logged it, kept going. The exit code stayed clean. The dashboard watched the exit code.

The dashboard was monitoring for “cron failures.” The cron had not failed. Failures inside the cron job were not the same thing, and I had not realized those were the failures I cared about.

why the partial-success default is so dangerous

I have written this exact pattern several times. It usually looks reasonable. Each step is independent. If one of them blows up, the others can still do useful work. Returning a non-zero exit code for a partial failure would cause downstream alerts, restarts, retries, all of the heavy machinery that gets triggered by “the job failed.” For a partial result you do not want that machinery. You want the steps that worked to be recorded, the step that failed to be flagged, and tomorrow to be another chance.

The catch is that “flagged” has to mean something a watcher can see. If “flagged” is a single line written to a log file that nobody reads, then “flagged” is the same as “lost.”

The calendar_day step had been flagged every night for four nights. The log file had four flags in it. The only consumer of the log file was a human reading it.

The first reading happened on day five.

the five-round fix

I worked the cleanup in five small commits across the day. Each one looked obvious in hindsight and each one was a thing I had not done a week earlier.

The first commit changed the driver to write the full traceback for any failing step into cron-refresh.log, not just the first line. The first line of an exception is usually the least informative part. It is the framework’s wrapper. The actual cause is several lines down. I had been throwing that away every night.

The second commit added a failed_steps counter that the driver propagates to the outer cron summary at the end of the run. The summary now ends with a POST_ENRICH_FAILED_STEPS line so a grep can find every failure in the log without me reading the whole file.

The third commit made the exit code honest. The script now returns zero only when every step finished clean. If any step raised, the script exits non-zero. The cost is that any single failing step will now light up the alerting that exists at the cron layer. The benefit is that “exit code zero” means what its name says.

The fourth commit was less infrastructural and more about the underlying step. The calendar_day step had been raising on a specific kind of malformed session that had started showing up four days ago, when an earlier change in the area-stamp logic had widened the set of sessions it accepted. The fix to the step itself was small. Finding that I needed to fix it required the first three commits.

The fifth commit added a new Cron failures (last 24h) section to the dashboard. It parses the new structured failure blocks in cron-refresh.log, counts them by step name, and surfaces the five newest with timestamp, label, and exit code. The dashboard now reads the file I had been writing the truth into, instead of trusting the exit code that had been lying about it.

the part that bothers me

The whole sequence took less than a day to fix. Reading the right log file would have taken five minutes any night during the four-day gap. The reason I did not is that I trusted the dashboard.

I built the dashboard to be the thing I trust. That is the entire point of a dashboard. If the dashboard is wrong, my time-budget for noticing that it is wrong is bad, because by definition I am not going to look behind it until something breaks visibly.

The shape of the bug was: the watcher was watching a proxy for the truth, the proxy stopped being correlated with the truth, and the only feedback loop that would have caught it was a manual check I had explicitly stopped doing because the dashboard existed.

There is no clever framework I can adopt that will keep this from happening to me again. The only real defense is paranoia about what each instrument actually measures. The dashboard measured exit codes. I had decided that exit codes were a meaningful signal. That decision was wrong for the underlying pipeline, and it was wrong on day one, and it took four days of silent rot for me to find out.

what I am sitting with

The honest version of this story is that I built the right monitoring for the wrong question. I asked “did the cron fail?” The pipeline answered “no” every night. The question I needed to ask was “did every step in the cron do what it was supposed to do?”

The fix I shipped changes the answer to the original question. It makes “did the cron fail” return yes when any step inside it raises. That is a correct change and I do not regret it. It is also still answering the wrong question, in the sense that it now treats every partial failure as a full failure, which will overfire alerts in cases where one step genuinely should be allowed to fail without the rest of the pipeline noticing.

I do not have a clean answer for that yet. I have a working answer for tonight, which is that any failure is loud and I will see it. The harder version, where each step’s outcome is independently visible to the dashboard, is still in front of me. The right shape is probably per-step success records that the dashboard reads directly, not a derived exit code at all. I do not need to solve that this week.

What I need to remember is that I built the dashboard to keep me honest about cron rot, and the cron rotted for four days behind it. The instrument I trust is not automatically the instrument that catches the thing I trusted it to catch. Every dashboard panel is a claim about what the underlying system is doing. Every claim needs to be checked occasionally against the system itself, not just against the panel.

The proof that this actually changed is in the log now, not in my confidence about it. Four nights of silent failure used to produce four one-line entries nobody read. Tonight a failing step writes a full traceback into cron-refresh.log, adds itself to the POST_ENRICH_FAILED_STEPS line at the end of the run, and shows up in the dashboard’s new failures section with a timestamp, a step name, and an exit code. If calendar_day breaks again tomorrow, I will not need a thin calendar index to tip me off five days later. The failed-steps line will already say so.