A forwarded GitHub notification landed in my inbox at 9:46 AM: semgrep had failed on our newest service repo, red X, no explanation beyond “1 check failed.” My first thought was that it was noise, some ruleset that didn’t like a new file glob. Six hours later it turned out to be neither noise nor a semgrep problem at all. It was our own deployment pipeline making an assumption about repo shape that had been true for years and stopped being true the moment we spun up a repo the GitHub-native way.

The forward that wasn’t the bug

That repo is one of a handful we’ve started building without the old scaffolding: no legacy branch layout, no manually wired staging and preview environments. Just main, PRs, and GitHub Actions doing what GitHub Actions does. When semgrep failed on it, I opened the run log expecting a lint complaint. Instead the job errored out trying to reference a staging branch that didn’t exist in that repo.

My first move was to patch it locally as a CI config problem. I edited the semgrep job to stop assuming a staging ref, pointed it at main, pushed, and watched the check go green. Good, done, forward the fix, move on. Except the actual failure hadn’t been in semgrep’s config at all, it had been triggered downstream, by our internal ticket-lifecycle tooling watching that CI run and trying to transition a ticket to “ready for staging” based on it. My patch made the semgrep job pass without touching the thing that was actually broken, so the next ticket that flowed through the same repo failed in the exact same place, just one step later in the chain. That cost me close to an hour before I stopped treating the CI log as the crime scene and started treating it as a symptom.

Following it up the chain

Once I stopped patching the job and started tracing what called it, the shape of the real bug came into focus. Our ticket flow has a fixed mental model of a repo: main branch, a staging branch, a preview branch, each one a rung the code climbs before it ships. Tickets move through the same three states in order, and each state transition is wired to check for the existence of the next rung. That model matches every legacy repo we have. It does not match a repo built GitHub-native, where there’s no staging branch and no preview branch, because the whole point of the newer setup is to skip that ladder and use PR-scoped environments instead.

The ticket tool wasn’t failing loudly. It was failing by assumption: it looked for a staging ref, didn’t find one, and quietly stalled the ticket in place rather than erroring in a way that pointed at the real cause. The semgrep failure was just the first place that assumption surfaced as something visible, because semgrep’s job definition happened to hardcode the same branch name.

What we actually shipped

The fix wasn’t a patch to that repo’s CI file, and it wasn’t a patch to semgrep. I walked it over to a teammate and we agreed on one flow for every repo, legacy or GitHub-native, instead of one flow that only worked for the repos we’d built five years ago. That meant treating “does this repo have a staging rung” and “does this repo have a preview rung” as things the tooling checks and adapts to, not things it assumes.

Out of that we wrote up a design doc and split the work into three tickets instead of trying to land it in one pass:

  • Build the missing rungs for the new repos, so that repo and its siblings actually get staging and preview environments that match what the legacy repos have. This one went first, since it’s the fastest way to make the new repos behave like the old ones without touching the tool itself.
  • Make the lifecycle tooling repo-aware, so it checks for a rung’s existence instead of assuming it, and adapts its transition logic per repo instead of hardcoding branch names.
  • Keep the original semgrep fix, but reclassify it. It’s not “the fix,” it’s the live use case that will prove the repo-aware tooling actually works once it ships.

That ordering matters. If I’d shipped the repo-aware tooling first without a concrete case that exercised it, I’d have been guessing at the interface. Keeping the semgrep failure alive as an unresolved case gave the second ticket something real to validate against instead of a hypothetical.

The part that generalizes

The actual lesson sitting underneath six hours and a wrong patch: a pipeline that’s only ever run against one repo topology doesn’t know it has an assumption baked in, and neither do you, until a second topology shows up and the failure surfaces somewhere unrelated to the actual cause. My first fix treated the symptom exactly where it appeared, in the CI job, and that’s precisely why it didn’t hold. The real question wasn’t “why did semgrep fail,” it was “what does our tooling assume is always true about a repo,” and that question doesn’t get answered by reading a failed check. It gets answered by reading everything upstream and downstream of that check until you find the place where an assumption was written down once, years ago, and never revisited.

We’ve got two new tickets in the queue now that exist purely to make our own tools less opinionated about repo shape. Neither one would exist if I’d stopped at the green checkmark.