I pushed a one-line fix to a deploy script, the deploy ran, and it failed at the exact line I’d just fixed. Not a different line, not a new error, the same line, with the same message I had spent the last ten minutes patching. For a second I was sure I’d botched the fix. I hadn’t. The fix was already on disk. The script just wasn’t running it yet, because it updates its own source in the middle of its own run and keeps executing the copy it loaded a commit ago.

If you have a self-hosted runner that pulls fresh code as part of the deploy, this is a trap waiting for you, and the symptom is a bad one: a correct fix that produces an identical failure. Here’s the whole thing.

The original failure

The job was “Deploy API” on GitHub Actions, and it went red in 37 seconds. The main API service itself started fine and passed its health check, so this wasn’t the app breaking. The job died later, in the deploy script’s service-management step, at the start helper, line 72:

Cannot find any service with service name 'worker-queue-api'.

The deploy script (deploy.ps1) walks a list of services, stops them, swaps the build, then starts them back up. That list is the main API service plus the two SMS-queue services. The two SMS-queue services aren’t actually provisioned on this box yet, so stopping and starting them is asking Windows for services that don’t exist.

The root cause was an asymmetry between the stop and start helpers. The stop helper guarded against missing services:

Terminal window
$svc = Get-Service -Name $name -ErrorAction SilentlyContinue
if (-not $svc) { Write-Warning "$name not registered (skipping stop)"; return }

The start helper did not. It went straight at the service and threw hard when it wasn’t there. So the deploy would happily skip stopping the two phantom SMS-queue services, then blow up trying to start them. The fix was obvious: mirror the same guard into the start helper so a not-yet-provisioned service logs a warning and is skipped instead of killing the run:

Terminal window
$svc = Get-Service -Name $name -ErrorAction SilentlyContinue
if (-not $svc) { Write-Warning "$name not registered (skipping start)"; return }

One file, one change. I committed the fix, pushed to main, and watched the run.

The twist

It failed again. Same job, same start helper, same Cannot find any service message at the same line. My change was sitting right there in main on GitHub, and the runner was producing output as if it had never seen it.

The instinct here is to assume you fixed the wrong thing. I almost re-opened the diff to hunt for a second copy of the bad code. Don’t. Look at the order of operations on a self-hosted runner that resets its own working tree.

Here’s what actually happens. The runner is itself driven by deploy.ps1. When the job starts, PowerShell parses the entire script into memory from whatever is currently on disk, which is the previous commit, because the tree hasn’t been updated yet this run. Then, partway through that already-parsed script, there’s a step that does:

Terminal window
git reset --hard origin/main

That command updates the file on disk to the new commit. But the script that’s running is the one PowerShell already parsed at the top, the old one. The start helper executing at line 72 is the old version, without the guard, even though the file on disk now has the guard. The script reset its own source out from under itself and kept running the version it loaded before the reset.

So the fix was in. It just wasn’t the one running.

Confirming it instead of assuming it

The tempting move here is “that must be it, it’ll pass next time” and a blind re-run. I didn’t want to re-run on a theory. I SSH’d into the box (the self-hosted runner, which runs on the prod machine) and read the on-disk deploy.ps1 directly. The guard was there. The git reset --hard origin/main from the failed run had already pulled the fix onto disk. So the on-disk script was correct, and the only reason the last run failed was that it had been parsed before that pull.

With the file confirmed, I re-ran the deploy. Green. The log showed exactly what I expected: not registered (skipping start) as a warning for each phantom SMS-queue service, and Health check PASS. The health probe only ever hit the main API service’s health check endpoint, so skipping the two missing SMS-queue services didn’t weaken what the deploy actually verifies. I checked the run output with gh run view --log.

Two runs, one fix. The first run failed because it was the run that installed the fix.

The real fix

Any deploy script that does git reset --hard origin/main (or any in-place self-update) partway through its own execution is running the previous version of itself on the run that performs the update. Interpreted languages parse the whole script up front, so the change you pushed lands on disk during the run but doesn’t take effect until the run after.

The cleanest fix is to kill the surprise entirely. Reorder the script so the self-update happens before anything load-bearing, then re-exec the freshly-pulled script in a child process:

Terminal window
# At the very top of deploy.ps1
if (-not $env:DEPLOY_REEXECED) {
git reset --hard origin/main
$env:DEPLOY_REEXECED = "1"
& $PSCommandPath @args
exit $LASTEXITCODE
}

& $PSCommandPath launches a fresh PowerShell process from the on-disk file, which is now the new version. The env flag prevents the loop. After this, the version executing and the version on disk are always the same commit, and a script can never fail just because it was the run that fixed itself.