Three verdicts, minutes apart, and none of them talked to each other

At 12:45 in the morning, the first review came back in three minutes: the funnel has no top. At 12:47, a second reviewer, working from the same document, took four and a half minutes and came back with a review dense enough to run tens of thousands of tokens on the same underlying problem. At 12:56, a third reviewer, given nothing but the document and no knowledge of what the other two had said, took almost fourteen minutes and wrote: “Do not lead with MCP or cross-vendor portability. Both are materially less validated than the strategy assumes.”

Three different AI systems. Same document. Same night. None of them saw each other’s answers. And all three landed on the identical flaw: we had built something on top of assumptions about how customers actually behave that nobody had ever checked with an actual customer.

The document was our positioning plan for this blog, dxdev.com. Not a memo. We’d built it out close to the real thing: the path a visitor would take, the case studies meant to prove the point, the pitch about why the underlying AI engine being swappable mattered to a reader. It looked and read like a finished plan, tangible enough that you could hand it to someone and they’d believe it was already decided. That’s the trap. A prototype that looks finished starts to feel true.

We had already spent real hours building it out before any of the three reviews landed. That’s the cost. Not money out the door, but the harder kind: time spent making a plan look convincing before checking whether the thing it was convincing people of was actually correct.

What “no top of funnel” actually meant

The first reviewer’s line, that the funnel has no top, sounds abstract until you translate it. The plan assumed that people read the blog first, then discover what we actually sell. Nobody had confirmed that path was real. We had built case studies to sit at the top of that funnel, and the reviewer’s point was that we were treating a guess about reader behavior as a settled fact, then building the entire rest of the pitch downstream of it.

The third reviewer’s point, about not leading with MCP or cross-vendor portability, was the same shape of problem from a different angle. We had a technical fact that sounded like a selling point: the system isn’t locked into one AI vendor, it can run on whichever one works best. Interesting to us. But nobody had confirmed that a customer weighing whether to hire us cared about that at all, let alone cared enough for it to be the headline.

Two separate technical claims, two separate reviewers, one shared root cause: we had confidence where we had no evidence.

Why running it three times at once mattered

If I’d sent that document to one reviewer, I’d have gotten one opinion and a plausible reason to argue with it. What made this different is that three systems, with no visibility into each other’s work, converged on the same underlying complaint using different words and different examples. That kind of agreement isn’t something you get by asking the same question twice and hoping for a different answer. It’s what happens when you stop asking “does this sound right” and start asking three independent readers “what’s wrong with this,” at the same time, and see if they land in the same place.

They did. So we stopped building the rest of the prototype that night. The plan had been to keep filling in the funnel, more case studies, more pages, more polish on the pitch. Instead we shelved it and went back to something slower and less satisfying: asking actual people, by hand, whether the assumptions in the plan were even in the neighborhood of true, before writing one more page that depended on them.

The prototype wasn’t wrong because it was badly built. It was wrong because it was persuasive before it was tested, and persuasive is exactly the quality that makes you stop checking.