The Same Red X, Three Days Running

I opened gh run list on a client’s API repo expecting one failure. I got a pattern. dependency-audit failed on 2026-08-04 at 09:24 UTC, on schedule, on main. It failed again on 2026-08-05 at 09:22 UTC, same workflow, same branch, same trigger. In between, the same workflow had gone green on pushes to develop and on a manual dispatch from a feature branch. So the picture wasn’t “CI is broken.” It was: the exact same problem is getting reported over and over, from slightly different angles, and nobody has looked at it long enough to name it.

That’s the part worth writing down, because the fix wasn’t a code fix. It was recognizing that the alarm was built on the wrong unit.

What the alarm actually looked like

dependency-audit (still filed as npm-audit.yml on disk, because branch protection rules match on filename and nobody wanted to touch that) runs on three triggers: every push to main, a release branch, and develop, plus a daily cron at 8:27 UTC so a freshly disclosed CVE pages someone even on a day with no pushes at all. That’s a reasonable design for catching new vulnerabilities fast. But it means the same unresolved issue gets to fail the build once per push per branch, and once per day forever, until someone closes it. The alarm fires per event. The actual problem lives per package.

And when it fires, what you get is a wall of table output. Package name, vulnerable range, patched range, a dependency path, an advisory link, repeated for every finding, dumped fresh on every single run. Nothing in that output says “this is the same thing you saw yesterday” or “here’s who owns this.” You have to read the Paths row yourself to find the actual culprit: apps__api > express-rate-limit > ip-address, vulnerable at <=10.3.0, patched at >=10.3.1, filed under GHSA-mwp4-54f8-5fhr. That’s the one fact that mattered, buried in a report that treats every run as a fresh incident instead of a status check on a known one.

Fixing the unit, not the noise

The actual fix wasn’t silencing the alarm. It was giving the unresolved issue a home that isn’t a CI run. The repo already had the mechanism: a documented exceptions list, docs/security-audit-exceptions.md, where an accepted CVE gets an owner, a review date, and a removal condition, because package.json doesn’t take comments and someone has to be able to say why a known-red audit is intentional. That’s the shift from “alarm per event” to “record per problem.” A CVE that’s been triaged and accepted stops being news every time a branch touches the audit step. It only becomes news again when the removal condition is met or the review date passes.

When the actual patch landed, the record proved its worth from the other direction. The commit that shipped it says exactly what happened in its own message: two new production highs, and a backport that made an existing exception obsolete. That’s the alarm doing its job in exactly one place. Not three branches independently telling you the same thing, not a daily cron re-announcing a fact that hadn’t changed since yesterday. One entry, updated once, when the underlying thing actually changed.

The lesson generalizes past dependency audits

Any alert system has an implicit unit: what counts as “one” thing worth telling you about. If that unit is smaller than the thing that’s actually broken, you get exactly this: the same failure multiplied across every branch, every push, every scheduled tick, dressed up as separate news each time. Three failing runs feels like three problems. It was one package, one advisory, one fix.

The other habit worth keeping is refusing to let a report just dump raw data and call that a message. A table of vulnerable versions is not the same as a sentence that says what’s actually wrong and who’s supposed to act on it. If I have to read the third column of a table to find the one fact that matters, the alert hasn’t told me anything yet, it’s just handed me the homework.

The GHSA-mwp4-54f8-5fhr advisory on express-rate-limit sat in that exceptions file for exactly as long as it took to matter and no longer: the moment a backport closed it, the fix commit said so in one line, and every future run of dependency-audit just goes quiet on that one line instead of relitigating it. That’s the whole difference between a report and a record, and it cost nothing but a markdown file nobody wanted to touch.