Why Most AI Pilots Stall Before Production
The failure is rarely technical. It is usually that nobody owns the output.
The standard explanation for a stalled AI pilot is that the technology was not ready. In the cases we have examined, that explanation is usually wrong, and the wrongness is consistent enough to be diagnostic.
A pilot typically succeeds on its own terms. It is run by a small, motivated group, on a well-chosen problem, with attentive humans in the loop and an implicit tolerance for rough edges. It produces a demonstration, a favourable metric and an internal reputation. Then it stops — not because it failed, but because the next step requires something nobody has been assigned to do.
That something is ownership of the output. In production, a system that generates a draft, a classification or a recommendation creates a stream of work that belongs to someone. Who reviews it. Who is accountable when it is wrong. Whose headcount absorbs the review burden. Whose budget pays for it once it leaves the innovation line. Whose process document must be rewritten, and who signs that rewrite. Pilots routinely skip all of this because the pilot team absorbs it informally, and the absorption is invisible in the pilot's reported cost.
The second recurring blocker is evaluation. A pilot is judged by whether the output looks good to the people who built it. Production requires a defensible answer to "how do you know", which means a labelled set, an agreed threshold, a monitoring regime and a decision about what happens when the number moves. Teams that did not build this during the pilot discover that retrofitting it costs more than the pilot did, and the request lands at exactly the moment enthusiasm is highest and patience lowest.
The third is data access under real conditions. Pilots run on an extract. Production runs against live systems with permissions, retention rules, regional constraints and a security review that did not participate in the pilot. It is common for a project to clear its technical hurdles in weeks and then spend two quarters on access approvals nobody scheduled.
None of these are model problems, and buying a better model does not address any of them. This is why the second pilot, run with a newer model after the first stalled, so often stalls at precisely the same point.
The teams that get through have usually made three unglamorous choices early. They name a business owner — not a technical owner — before writing code, someone whose existing metrics improve if the system works and who has authority over the process it touches. They build the evaluation set during the pilot rather than after, accepting a slower demonstration in exchange for a defensible one. And they scope the first production deployment to a single workflow narrow enough that the review burden can be measured honestly, then widen only when that measurement is boring.
There are limitations to this account. It is drawn from deployments that were discussed openly, which biases towards organisations willing to talk about failure and away from those where the pilot was quietly absorbed into a roadmap. Published survey figures on pilot-to-production rates vary widely and use inconsistent definitions of both terms; we would not put weight on any specific percentage. And some pilots genuinely do fail on capability — verification-heavy work in regulated settings remains hard — though in our sample that is the minority case rather than the default.
The practical implication is uncomfortable for the way most organisations structure this work. Innovation functions are optimised to produce pilots, and a pilot is a cheap artefact to produce. Production requires a transfer of responsibility into a line function that did not commission the work and is not rewarded for adopting it. Until that transfer is designed deliberately — with named owners, agreed measurement and honestly costed review — the pipeline will keep producing successful demonstrations that go nowhere, and the post-mortem will keep blaming the model.
Every claim above is sourced to a document, a named person, or a record we hold. Where we could not verify a claim, we say so. Read our standards and corrections policy →
