The Railway build log read:
ERR_PNPM_OUTDATED_LOCKFILE Cannot install with "frozen-lockfile" because pnpm-lock.yaml is not up to date with package.jsonCloudflare Pages had deployed the exact same commit ten minutes earlier without complaint. Same repo, same dependency bump, one green and one red. That gap is where I learned that “the agent’s memory” is a bigger surface than the chat context window, and that most of it fails silently.
Two lockfiles, one dependency tree
I’d bumped a shared utility package used by both a Cloudflare Pages frontend and a Railway backend. I tested locally with npm install, ran the CF Pages build command (npm ci) to confirm it worked, and pushed. It did work, because npm ci reads package-lock.json, which I’d just regenerated. Railway’s build uses pnpm --frozen-lockfile, which reads pnpm-lock.yaml, a completely separate resolved-dependency file that I hadn’t touched.
The assumption I acted on, and paid for, was that a monorepo has one memory of its dependency tree. It doesn’t. It has one memory per package manager, and updating one doesn’t touch the other. npm ci and pnpm --frozen-lockfile both fail loudly if their own lockfile is stale, but nothing fails if only one of two divergent lockfiles gets refreshed, because each build only ever checks its own copy. The frozen-lockfile flag is exactly why this is --frozen-lockfile and not --force, and I still burned a deploy cycle chasing an error that had nothing to do with the actual code change.
What I do now: any time a dependency changes, I run both npm install and pnpm install in the same commit, regardless of which platform I’m actually working against that day. Two lockfiles, one commit, no exceptions. It’s not elegant, it’s a checklist item, but it converts a silent divergence into a diff you can see before you push.
The classifier that forgot itself twice in one session
The same class of bug showed up somewhere with no lockfile at all. I have a memory system that records lessons across sessions, meant to stop the agent from repeating mistakes. One recurring mistake was misclassifying which project a piece of work belonged to, treating a personal side-project fix as platform work, or the reverse, and filing the ticket in the wrong queue.
The first time it happened in a session, I had it write a memory note: classify the work’s domain before creating any ticket. Later in that same session, it happened again. Not in the next session, the same one, after the correction had already been recorded. That’s the detail that mattered: this wasn’t a failure of long-term retention across sessions, it was a failure to apply a rule it had just written down, on the very next relevant decision, minutes later.
That told me the fix couldn’t live in memory at all. Memory recall is retrieval, it depends on the current prompt surfacing the right note at the right moment, and ticket creation wasn’t reliably triggering that lookup. So I moved the check out of memory and into a mechanical gate: a domain-classification step that runs before every ticket is created, every time, not something the agent has to remember to consult. The rule went from “something it should recall” to “something it cannot skip.”
DNS that remembers the wrong thing
The third version of this showed up on Cloudflare. I was setting up CF-for-SaaS cross-account hostname validation, and the custom hostname record was in place, resolving correctly under dig. Validation still wouldn’t complete. The record existed, but it was DNS-only, the grey cloud in the dashboard rather than the orange one. For CF-for-SaaS, the origin has to be proxied for the validation handshake to work at all. An unproxied record isn’t a broken record, it just isn’t the kind of record this feature reads. Nothing in the DNS tooling flags that distinction as an error, because as far as plain DNS is concerned, the record is fine.
Now when a custom hostname won’t validate, checking whether the record is proxied is the first thing I look at, before touching certificates or account permissions, because the failure mode looks identical to half a dozen other causes and only one of them is a single toggle.
What these three have in common
None of these are context-window problems. They’re all long-lived state that something in the pipeline treats as authoritative: a lockfile, a memory note, a DNS proxy flag. And in every case, the system that reads that state doesn’t verify it against reality, it trusts it. pnpm --frozen-lockfile trusts the lockfile is current. Ticket creation trusted a note would get recalled. CF-for-SaaS trusted the orange cloud was set. When any of those go stale, the failure shows up somewhere else entirely, in a build log, in the wrong project queue, in a validation status that never flips, and by the time you see it, you’re debugging the symptom instead of the actual drift.
The fix isn’t a smarter agent. It’s fewer places where “remembered” and “true” are allowed to diverge without a check that catches it.