The load average sat at 4.2 on a box that should idle at 0.3. ps aux showed six next dev processes, oldest one three days alive, none of them attached to a terminal anyone was looking at. Somewhere in the middle of that list was a find / that had been walking the filesystem for eleven hours, kicked off by a script that grepped for a config file and never bothered to scope the search.
That was the actual state of the machine when I went looking for why everything felt slow. Not a crash, not an alert, just a system quietly losing capacity to processes nobody remembered starting.
Why the servers piled up
The dev-server launcher had no memory of what it had already started. Every session that needed a preview would just run npm run dev (or the equivalent for whichever repo it was in) and bind to whatever port was free. If port 3010 was taken, it grabbed 3011. Next session, same thing, another port. Sessions end all the time in this setup, some cleanly, some because the agent got killed or the terminal closed without a shutdown signal. The server keeps running either way. Nothing in the lifecycle ever asked “is this already up” or “should this still be up.”
Six months of that and you get six orphaned Next.js dev servers, each holding a chunk of memory and doing hot-reload file watching on a repo nobody’s editing anymore, plus a background find some earlier session had left crawling the disk for a lockfile it should have looked for with a targeted path in the first place.
The wrong turn: killing it by hand
My first move was the obvious one: write a one-off script to pkill -f "next dev" and clear the find. It worked, load average dropped back to normal in about ten seconds. I closed the ticket in my head and moved on.
It came back within a day. Of course it did, because I hadn’t touched the thing that was creating orphans in the first place, I’d just swept the floor once. The launcher still had no idea a previous instance existed, so the next session spun up its own dev server on the next free port, and the process that had already been orphaned an hour earlier was still sitting there because nothing had told it to die. Manual cleanup treats the symptom every time you remember to run it and does nothing the rest of the time. That’s the cost of the wrong turn: a day where I believed the problem was solved and it wasn’t, and the leak kept compounding quietly under a load average that looked fine right after I checked it.
What actually fixed it
Two pieces, tracked under a single internal ticket.
An idempotent dev-server launcher. Before starting anything, it checks for an existing process bound to the port or matching the repo path, by PID file, not by a fragile process-name grep. If one’s already running and healthy, it reuses it and returns immediately instead of spawning a duplicate. If the PID file points at a dead process, it cleans the stale record and starts fresh. Same command, run five times in a row, produces one server, not five.
The reaper. A script that walks known process patterns (dev servers, long-running find/grep invocations past a runtime threshold, orphaned bridge spawns) and kills anything that no longer has a live session behind it. The key design choice was where it runs, not as a cron job hoping to catch things eventually, but wired directly into session start and session close. On start, it sweeps anything orphaned by the previous session before doing any work. On close, it sweeps anything the session itself spun up that’s still alive. That closes both ends: you can’t accumulate garbage between sessions, and you can’t leave garbage behind when a session ends, clean or not.
Together they mean the launcher won’t create a duplicate and the reaper cleans up the ones that already existed before this shipped, plus anything that slips through some other way (a hard kill, a crashed terminal, an orphaned bridge spawn from the handoff path). Committed and pushed the same session, verified against the running box afterward: load average back to 0.3, ps aux showing exactly the processes that had live sessions behind them, nothing else.
Why session boundaries, not a timer
I could have made the reaper a cron job running every fifteen minutes and called it done. I didn’t, because a timer-based sweep still lets orphans live for however long the interval is, and it runs whether or not anything actually needs sweeping. Hooking it to session start and close means the check happens exactly when the conditions that create orphans (a session ending, a new one beginning) are actually occurring. It’s a smaller number of checks, each one meaningful, instead of a constant poll hoping to catch a leak in the act.