The full test suite takes 497 seconds. For as long as anyone had written it down, the number in our notes was 26 minutes, and nobody had ever timed it. That wrong number turned out to be the thread that led to a guard test that had never executed once.
The test that was supposed to protect the hotfix rules
We have a hook that enforces our hotfix rules on hotfix branches. Beside it sits a pytest file with two kinds of assertions. Deny cases feed the hook a forbidden action and expect a refusal. One allow case feeds it a legitimate action and expects a pass. The allow case is the one that proves the guard isn’t just refusing everything.
The allow case was green, and had been for as long as anyone looked.
I found out why while moving that same file off a live working clone. The old test ran against a real checkout on disk, so I was rewriting it to build a throwaway repo in tmp_path and adding a conftest autouse fixture that fails any test that mutates git outside tmp_path. To write the fixture I had to read exactly how the test launched the hook, and the answer was one line:
bash = shutil.which("bash")On this Windows machine, shutil.which("bash") returns WSL’s bash, because System32 sits ahead of Git’s bin on PATH. The hook is written for git-bash, with its path conventions and its /c/... style mount points. Under WSL bash the script bailed on its first path lookup and exited 0 without evaluating anything.
Why a bailed hook reads as a pass
The allow assertion was roughly “run the hook on a legitimate action, assert it does not deny.” A hook that dies quietly on line one does not deny anything. The assertion held, the test went green, and the guard was never reached.
The allow case is the positive control, and it was the one that could pass by doing nothing. I also never worked out why the deny cases didn’t flag the dead hook, which I’d have expected them to.
The wrong turn: trusting green output I couldn’t read
Earlier that same day I had pointed a verification agent at a different new module and told it not to trust the builder’s “12 passed, exit 0.” Our tooling compresses pytest output, so the raw result comes back as a single line like Pytest: 12 passed. The agent tried to get past the compression by redirecting output to a file. The first attempt wrote to /tmp_out.txt, which under git-bash resolves to a protected directory:
/usr/bin/bash: /tmp_out.txt: Permission deniedEXITCODE=1That exit code was the redirect failing, and it told us nothing about pytest. The agent moved the target into the scratchpad and got EXITCODE=0, but the file it read back still held just the one compressed line.
My own instruction to the agent had a subtler defect. python -m pytest ...; echo "EXIT=$?" is fine, but the same idea written as pytest ... | tail -60; echo "EXIT=$?" reports the exit status of tail, not pytest. The first run in that session printed EXIT=0 this way. It looked like confirmation and confirmed nothing. We had spent a chunk of a session collecting “proof” that was structurally incapable of failing.
We treated a compressed, piped, zero-exit result as verification, and that is exactly the trust that let the bash test sit green. Same failure, two layers.
The fix
Three changes.
Pin the interpreter. The test now resolves git-bash explicitly from its install location and skips loudly if it is missing, instead of falling back to whatever PATH offers. A skip shows up in the report. A vacuous pass does not.
Give the allow case a witness. The hook now has to prove it ran. The test asserts the hook produced its expected marker on stdout before it checks the decision. A hook that bails yields no marker, and the test fails at that assertion instead of passing.
Prove the test can fail. I broke the hook on purpose, replacing the rule check with exit 1, and confirmed the allow case went red. Then I put the pinned bash back and confirmed it went green for the right reason. We should have done this the day the test was written.
The 26 minutes
While I was in there I timed the suite for the first time. The full run is 497 seconds. The 26 minutes in our notes had never been measured. A suite people believe takes 26 minutes is a suite they run rarely, and a test that rarely runs can be wrong for a long time.
So I added a fast subset that skips the slow integration files. It runs 2,859 tests in 71 seconds, short enough to run before every commit. Writing that number down as measured, with the date, was as much a part of the fix as the bash pin.
What we check now
A green test that does not reach its target looks identical to one that does. Three cheap habits close most of that gap:
- Every allow or positive-control test asserts a marker that only the real code path can produce.
- Any test that shells out resolves its interpreter explicitly.
shutil.whichon a machine with WSL installed answers with the wrong shell. - When we want an exit code, we capture it from the command itself, with no pipe between them.
The guard test now fails when the hook doesn’t run. It took a timing exercise on an unrelated number to make me open the file.