78 of 12,866 requests that day came back as IIS’s own 502 page, on a Chrome upload endpoint’s transcript replay. I hit it at 17:44 UTC on my phone: the cockpit-server was running under tsx-watch, and every file save restarted Node mid-request. We moved it to the built production server, node dist/index.js with NODE_ENV=production on port 3010, stopped the spine loader from painting the raw upstream error body over the UI, added a client-side self-retry, and made the server refuse to start at all if PORT wasn’t set explicitly, instead of falling back to a default and 502ing forever behind the proxy.

That same day we shipped RichInput.tsx, a single composer with type, mic, attach, paste and drag-drop, wired into every input surface on the cockpit: gate cards, the capture rail, the closure queue, the task chat panel. Wiring in the mic button is what surfaced the real problem. Voice transcription had been failing for weeks, ever since someone pulled the OpenAI key out of the cockpit-server’s service environment to close an earlier issue. The reasoning at the time was sound: the same Windows service also runs a scheduled classification worker, and a key sitting in process.env would get handed to that unattended worker along with everything else the service does. So the key came out, and voice quietly started failing closed. Nobody noticed because nothing in the app was calling it yet.

The fix mirrors a pattern already in place for speakTts: server/pwa/voice.ts now reads the key from disk, inside the request handler, and never assigns it to process.env. That’s the part that’s easy to get backwards. My first instinct was to just re-add OPENAI_API_KEY to the NSSM service env and move on, since the endpoint clearly needed it and the env var was right there in .secrets/openai.env. I actually did this, restarted the service, and transcription worked. Then I remembered why it had been pulled: the classification worker runs on the same process and reads the same env, unauthenticated, on a schedule. Re-adding the key didn’t just fix voice, it silently re-armed a background job that could burn OpenAI spend with zero user in the loop and zero visibility if it went sideways. I pulled the key back out within the same session, before the scheduled worker’s next tick, and started over from disk-read instead of env-var.

Reading from disk solved the leak path but not the actual risk, which was that a live-mic endpoint plus a valid key is an open spend spigot with no guardrail. So the rewrite added three things past the bare fs.readFile:

  • A vault-key gate: the read only succeeds if the caller is authenticated against the same session-key check the rest of the bridge API uses.
  • A test-runner refusal: the handler checks for the CI/test environment marker and returns an error instead of hitting the network, so a stray test run can’t spend real money.
  • A 1 hour/day audio cap per session, plus a visible spend ledger the transcription write path appends to on every call, so the number is never a mystery days later.

Same day, POST /api/upload went in for boxes that don’t have a session yet, saving to <cwd>/.cockpit-uploads/<sid>/<ts>-<safeName> with the filename stripped of path separators, nulls, and control characters, then clamped to 180 characters. Verified live before calling it done: transcription returns 200 with a ledger entry written, upload returns 403 unauthenticated and 200 authenticated with the file actually landing on disk, and the full suite ran 129 green.

The lesson that stuck wasn’t “add a cap.” It was that pulling a key to close one hole (the scheduled worker) created a second, invisible hole (a dead feature nobody was watching), and the only reason we caught it was that an unrelated UI pass happened to touch the same wire. A cap and a ledger don’t just stop runaway spend, they make it visible enough that the next time we pull a key for one reason, we’ll actually notice what else went dark.