I was building an autonomous PM loop, the kind of thing that wakes up on its own, looks at what is open, and proposes the next move. The hardest decision of the day was not a line of code. It was one question I had been quietly avoiding: where does the thing that keeps waking up actually run?
That question is the real fork in any always-on agent design, and it is easy to skip past it straight to a daemon because a daemon is what everyone reaches for. I left the fork open in the design doc on purpose, because I did not trust my first instinct, and ran it past codex before committing.
The two shapes
There are two honest answers to “where does the loop live.”
The first is a resident daemon. A long-lived process that sits in memory, holds the orchestration state, and ticks on its own timer. This is the default mental model. It is also the one that quietly signs you up for a pile of new responsibilities: it has to survive reboots, it needs a supervisor to restart it when it dies, it owns its own state in memory that vanishes if it crashes mid-cycle, and on Windows it means standing up another service to babysit.
The second is a periodic tick. Something external wakes a short-lived process every N minutes, it does its job, it writes its state to disk, and it exits. The orchestration state does not live in a process at all. It lives in files. The “daemon” is just whatever fires the tick.
The reason this matters: restart-safety is free in the second model and expensive in the first. If your state is on disk and your worker is stateless and short-lived, a reboot is a non-event. The next tick picks up exactly where the last one left off because there is nothing in memory to lose.
What codex called, and why it was right
Codex’s recommendation was the second shape, with one specific twist that made it click.
Use a scheduled-task tick (a Windows schtask, the cron equivalent) for the orchestration state. But do not have that tick spawn agents directly. Dispatch approved work through the existing in-process endpoint the server already has: POST /api/bridge/spawn, backed by bridgeRegistry.spawnSession({cwd, seed, model}).
The “twist” is the load-bearing part. The server already owns a gnarly piece of platform plumbing: spawning claude.exe under the right Windows context. It runs as LocalSystem (NSSM default), where USERPROFILE points at the wrong place, so every claude --resume fails with “No conversation found” until you override the home directory via a CLAUDE_USER_HOME env var. That code is hard-won and it lives in the server process. If the tick spawned agents itself, it would have to re-implement that resume dance from a non-interactive scheduled-task context, which is exactly the environment where Windows logon types and PATH resolution and drive mapping all bite you. By routing dispatch back through the server’s existing spawn endpoint, the server keeps owning the workaround. The tick never has to know about it.
So the responsibilities split cleanly:
- The tick: resolve scope, classify what is open, write a proposal to
.runs/pm/. That is all. It is not a dispatcher. It is a classifier that leaves a note. - The server: when a human approves a proposal, the gate-check runs and
spawnSessionhappens in-process, right where the resume workaround already lives.
That separation is what makes the whole thing restart-safe by construction. The tick is stateless and short-lived, so it cannot leak a half-finished cycle across a reboot. The state is files in .runs/pm/ plus the ticket’s plan checklist. And the dispatch path is the same in-process endpoint the server already owns, so there is no second code path to keep in sync. How the human actually signals approval (a file flag in .runs/pm/, a button in the server’s own UI, or a CLI command) I left open in the ADR. That control surface gets decided before any auto-dispatch-on-approval gets wired.
There is one real risk this introduces, and codex flagged it: tick overlap. If a tick runs long and the next one fires before it finishes, you get two workers stepping on the same state. The fix is a lease or lock in .runs/pm/, the standard guard for any cron-style job that must not run concurrently with itself. Worth building before you trust the loop, not after.
What I did not have to build
This is the part that justifies the whole detour. By taking the tick shape, I reused a mechanism I already had.
A nightly refresh job already runs as a Windows scheduled task. The team already had the muscle memory and the tooling for “a schtask fires a short Python/Node job on a timer.” Adding a new task that fires every 30 minutes was not new infrastructure. It was a new entry in a pattern that was already load-bearing and already monitored.
So the ledger came out like this. New always-on infra: none. New service to supervise: none. New restart-recovery logic: none, because the state is on disk and the worker is stateless. New elevation handling: none, because the server still owns it. What I actually added was a scheduled-task entry and a classifier that writes a file.
Compare that to the daemon path: a new long-lived process, a service wrapper to keep it alive across reboots, in-memory state that needs its own persistence story, and a re-implementation of the LocalSystem resume workaround in a context where it is hardest to get right. The tick shape erased all four.
Verify the seam before you build on it
One more thing, because this is where I almost got burned. Codex gave me a confident answer and a line citation for where the /api/bridge/spawn seam lived. The citation was off. The file is 1218 lines and the line it pointed at was not the endpoint.
This is the failure mode with any LLM recommendation that names a specific location: the reasoning can be exactly right while the citation is exactly wrong, and the wrong citation is the part that looks the most authoritative. If I had taken “it’s at line X” on faith and the seam had not actually existed in the shape I assumed, the entire restart-safe story would have been built on air.
So I opened the file and confirmed app.post("/api/bridge/spawn", ...) actually existed and that bridgeRegistry.spawnSession actually took the {cwd, seed, model} payload I was going to dispatch through. Once verified, I captured the whole call as a short internal design-decision note, so the next person (or the next me) does not have to re-derive why the loop lives in a schtask instead of a daemon.
The decision survives because the seam is real, not because codex said so.
The takeaway
Before you build an always-on daemon, ask one question first: does an existing periodic-tick mechanism, plus an existing in-process dispatch endpoint, already give you restart-safety for free?
If your orchestration state can live in files and your worker can be short-lived and stateless, a tick beats a daemon on almost every axis that matters in production. No new service to supervise, no in-memory state to lose on reboot, and the hard platform workarounds stay where they already work instead of getting copied into a worse context. The tick becomes a classifier that leaves a note, and the dispatch stays on the path a human already trusts.
And when an LLM hands you a recommendation with a line citation, open the file. The reasoning can be sound and the citation can still be wrong, and the citation is the part you are about to build on.
Related
- The console window that flashed every 30 minutes: Interactive vs S4U scheduled tasks: the scheduling context difference that determines whether a task runs visibly or silently
- claude —resume worked in my terminal and failed in my web app: the LocalSystem homedir trap: the LocalSystem/USERPROFILE problem this architecture routes around
- Cloudflare Tunnel as a Windows scheduled task (and the visible-console bug I shipped): another case of reusing the scheduled-task tick pattern on the same box
- You can’t repoint a junction under live file handles (why I had to stage the .claude move): Windows-specific constraint that shapes how services and agents coexist
- My pinned app kept vanishing after every reboot, and it wasn’t Windows being flaky: restart-safety failure mode that the tick-over-daemon choice avoids