Evaluation Is Quietly Becoming the Product
As models converge on capability, the ability to prove a system works is the remaining differentiator.
When several models perform within a few points of each other on the public benchmarks, the benchmark stops functioning as a purchasing signal. What replaces it is harder to buy, harder to fake, and almost never included in a vendor demonstration.
Teams that have put AI into consequential workflows tend to have built the same thing, independently and without much fanfare: a private evaluation set drawn from their own data, scored against outcomes they actually care about, run automatically on every change to the prompt, the retrieval layer or the model. It is a few hundred to a few thousand examples, curated slowly, argued over, and treated as an asset.
The reason this matters more than model selection is portability. A company with a good evaluation harness can change providers in an afternoon, because it can measure whether the swap made anything worse. A company without one cannot change providers at all in any responsible sense; it can only change them and hope. Model loyalty, in other words, is usually a symptom of missing measurement rather than a considered technical position.
Building the harness is less about tooling than about deciding what counts as right. This is where the work actually is, and it is editorial rather than engineering work. For a support assistant, is a correct-but-unhelpful answer a pass. For a contract reviewer, is a missed clause equivalent to a false alarm, or ten times worse. For a summarisation task, who adjudicates when two reviewers disagree. Teams that skip this arrive at a number that moves without anyone knowing what it means.
A few patterns recur among harnesses that survive contact with production. The examples come from real traffic, including the awkward cases, rather than from synthetic generation — synthetic sets tend to encode the assumptions of whoever wrote the prompt. Scoring mixes automatic checks with periodic human audit, because model-graded evaluation drifts and needs calibrating against people. The set grows from incidents: every production failure becomes a test case, which is how the harness stays aligned with what actually goes wrong. And the threshold for shipping is agreed in advance, in writing, by someone accountable for the outcome rather than for the release.
The strategic consequence extends beyond engineering. Evaluation capability determines how fast an organisation can adopt anything new. When a materially better or cheaper model appears — and on current cadence one appears every few months — the constraint on capturing that improvement is not integration effort but the ability to prove nothing broke. Organisations that can run that proof in a day compound the industry's progress. Those that cannot are effectively frozen at whatever they deployed first.
The limitations of this position deserve stating. A private evaluation set is only as good as its coverage; it measures the failures you have already imagined or already suffered, and is systematically blind to novel ones. Model-graded scoring shares failure modes with the systems it grades, and can be gamed by prompt changes that improve the score without improving the output. Small evaluation sets produce noisy comparisons, and a two-point difference on three hundred examples is frequently nothing at all. And none of this substitutes for monitoring live behaviour, where the distribution of real inputs shifts in ways no fixed set captures.
With those caveats, the direction is clear enough. As capability converges, the differentiator is not which model an organisation uses but whether it can demonstrate that the system built on it works. That demonstration is becoming the product — internally, in procurement, and increasingly in front of regulators who are less interested in the model card than in the evidence.
Every claim above is sourced to a document, a named person, or a record we hold. Where we could not verify a claim, we say so. Read our standards and corrections policy →
