The kernel’s v3 reached Railway MySQL production on May 14 with 201 passing tests, 5 tRPC procedures, a publish CLI, a SPA RunView, and a working demo flow. I could paste a URL into a browser and watch the full onboarding loop run end to end. That felt like the end of something.
the kernel was working
The architecture had been through five review rounds. Claude had audited the tRPC contract twice. The database schema lived on Railway with a proper migration chain. The publish CLI serialized skill snapshots to a JSON wire format and the RunView consumed them on the other end. None of this was prototype code. It was small and clean and deliberate.
When I stood it up in the real environment for the first time, two bugs surfaced that no amount of review would have caught.
The first: SHA-prefixed hash identifiers in the wire format were too long for the VARCHAR(64) columns in the Railway schema. The schema had been designed against an older format. The hash shape changed in a late refactor and nobody noticed because the test fixtures used short stub IDs. The fix was four characters in a migration file.
The second was subtler. MySQL silently reorders JSON object keys on write. The integrity check downstream was computing an HMAC over the raw JSON string. Fetch the value back out of MySQL and the key order is different. The HMAC fails. Everything downstream of that check reports corrupt. This one took two hours because the error message said “integrity violation” and not “key reordering,” and the first three places I looked had nothing to do with JSON serialization.
Fixed both. Pushed. Tests green. I could have closed the ticket and called it a good day. The kernel was objectively working end to end in a real environment.
what a colleague’s house surfaced
I tested it with a non-technical user. The session lasted roughly ten minutes. By the end of it I had a clear answer to a question I had not thought to ask.
The flow I had built required the user to read a step, navigate to the thing the step described, perform an action, copy a value out of that action, come back to the onboarding UI, paste the value in, and repeat. I called this the paste-back-everywhere model in my notes afterward. At the time I thought of it as a reasonable tradeoff between implementation complexity and user autonomy.
The user did not experience it as a reasonable tradeoff. He experienced it as worse than the static HTML doc he had been using before. The old doc was a single page he could read at his own pace. The new system interrupted him at every step to ask for a value he had to go fetch from somewhere else.
He was not wrong. The interaction model was wrong.
archaeological layers
I went back through the architecture with that observation in mind. The v3 design, and the v4 and v4.v2 iterations I had roughed in during the same sprint, all shared the same assumption: the user is the thing that moves information between systems. They carry values from their cloud console to our onboarding UI. They paste. They confirm. They advance.
That assumption was legible in the data model, in the tRPC contract, in the RunView component. The whole architecture was built to support a human acting as a shuttle.
That was the wrong premise. Five review rounds and 201 tests had sharpened the implementation of a bad premise. The review rounds were real work. They caught real issues. But they were asking “is this correct” about a thing answering the wrong question.
v3, v4, and v4.v2 are not the destination now. They are the layers you can see when you cut through the hillside. They show the thinking. They are not the product.
what the product actually is
The onboarding platform needs a different shape. Skill is the atom. A skill is a bounded, executable chunk of setup: here is the thing, here is the state before, here is the verification that it worked. Devices are first-class. The same skill runs differently on a Windows desktop versus a Linux server versus a hosted environment, and the system knows that before it sends the first step.
Claude fetches the next step. The user does not shuttle anything. The agent reads the current device state, determines what the next action is, and either performs it or surfaces exactly one confirmation request. The difference between v3 and this model is the direction of information flow.
In v3, the user pushed values into the system. In the new model, the system pulls state from the environment and the user approves actions. That is a different product with the same name.
The database schema, the skill snapshot format, and the tRPC contract all survive in some form. The RunView does not. The publish CLI survives in a different role. The 201 tests describe something that will be refactored rather than discarded.
the week was not wasted
I have heard the argument that shipping a wrong architecture is just waste. I do not think that is right.
I know the exact failure mode of the paste-back-everywhere model because I built it fully and watched a real user bounce off it. I know the MySQL JSON key-ordering trap because I hit it with real data on Railway hardware. I know the SHA column width issue because the migration ran against actual schema, not a fixture stub.
None of that knowledge came from the five review rounds. It came from ten minutes in a colleague’s house and two hours on the HMAC failure.
A week of clean architecture bought the ability to recognize a wrong premise quickly. The wrong premise did not survive contact with one non-technical user. Elegant architecture in service of a bad interaction model is a week of tuition, not a week of waste. The next version starts from the right question.
Shipping the kernel does not mean you found the product. Real users invalidate elegant architecture in minutes, and the minutes are the data you needed.