The number that went missing

One of our own published posts carried “working-queue-grew-to-65” in its slug and never once printed 65 in the body. Its own verify_notes admitted the omission: it had cut “ticket counts, field behavior, and internal workflow detail.” Six consecutive posts before it shipped with the identical three-heading skeleton, and two of them closed on the same sentence, word for word. That’s what an automated content pipeline looks like when nobody audits its output for a while.

I wrote the rubric that told it to do this. The instruction to “redact internal architecture specifics” and “generalize” unsupported numbers was mine, added with anonymization in mind, and it shipped six consecutive posts before I looked at what it had actually been publishing under my own review process. This is exactly the blog whose credibility depends on receipts over narrated conclusions, running a pipeline that was quietly manufacturing narrated conclusions in place of the receipts I’d told it to delete.

We run a daily sweep that mines the work log and the JIRA timeline for blog seeds, hands each seed to a drafter, runs the draft through a review rubric, and promotes what passes. Two separate defects were producing that slop, and they were both load-bearing enough to file as root causes rather than symptoms.

Root cause one: the rubric instructed deletion

The review rubric’s own text told the editor pass to “redact internal architecture specifics.” For a build-log post, architecture specifics are the entire point. The same rubric told it to replace a real business with “a client” and to “generalize” any number the draft couldn’t fully support. Generalizing a number doesn’t anonymize it, it just deletes the fact and leaves a vague sentence standing where a real one used to be.

The fix rewrites both the drafting prompt and the editor rubric around one distinction: substitution, never deletion. Names, company domains, and ticket ids get replaced with an equivalent-shaped stand-in, “a 63-team league” instead of “a client.” But numbers, thresholds, durations, error strings, commands, field names, and tier names are explicitly listed as things that identify nobody and must survive verbatim. And for a claim the sources genuinely can’t back, the rule flipped from “soften it” to “cut it entirely.” A missing sentence is honest. A hedged one just wears the shape of a fact it can’t stand behind.

Root cause two: the drafter had nothing to draft from

The second defect was upstream. Nineteen backlog rows cited a repo @sha and resolved to zero sources, because the git resolver only ever checked one repo root, and one of the archived repos lived on a different drive. When a resolver silently returns nothing, the drafter doesn’t refuse, it fills the gap with something plausible. That’s where invented specifics come from: not malice, just an empty context window getting padded.

A second version of the same failure showed up on range citations. One row cited a twelve-year commit range; git log --stat over that range emitted 8.8MB and blew the resolver’s 30-second timeout, so the row fell through exactly like the missing-repo rows did. The fix summarizes instead of dumping: commit count, the oldest three commits in the range, the most recent forty subjects. That fits the token budget and is still true. We also added resolution for two citation shapes the original resolver never covered: <repo> <TICKET> commits, resolved by grepping the log for the ticket key, and shorthand ticket runs like 7558/7560/7566/7570-7572, which previously only yielded the first number in the string.

The alternative we didn’t take

A related module, the identity scrub gate that runs after the drafter, faced a version of the same choice: eight literal banned tokens, warn-only, so a leak could still ship. The obvious hardening is to make that guard fail closed. We measured what that would actually do first: 71% of drafts in the queue would start failing, which means the daily cron produces nothing and nobody notices, because a silent empty output doesn’t page anybody the way a leak does. Trading a leak for a silent stall isn’t a fix, it’s a different failure mode wearing a safer costume. What we built instead scrubs by substitution first, using a deterministic per-identifier stand-in with a private mapping file kept alongside the draft, then re-scans the result and only fails closed if something survives the substitution pass. The pipeline keeps moving on the common case; it stops hard only when the automated fix genuinely didn’t work.

Fifteen new tests pin the rubric and resolver changes; the full suite ran 416 passed against 7 pre-existing failures unrelated to this work. Two gaps came out of the audit that aren’t closed yet, filed as their own tickets rather than folded into this fix, along with a stray false solo-founder claim sitting in the generator prompts that’s built and verified on a branch but still waiting to merge. Anonymization was never the thing eating these posts. Treating “redact” and “delete” as synonyms was.