Ask a company running a generative AI system in production how it knows the system is working, and the answers cluster into a revealing pattern. Users report satisfaction. Volume is up. The team has not received complaints. Occasionally someone mentions a benchmark the vendor published.
What is usually absent is a measurement of whether the system does its actual job better than the process it replaced, evaluated against cases where the right answer is known.
The pilot metric that did not survive contact
This is not carelessness so much as a structural feature of how these deployments happen. A pilot is scoped narrowly and measured, because it has to justify expansion. The expansion inherits the credibility of the pilot without inheriting its measurement, since applying the same rigor at production scale requires labeled data, sustained human review and someone whose job it is.
The difficulty is genuine. Evaluating a summarization, a drafted response or an extracted field requires knowing what the correct output was, and for most business processes nobody ever wrote it down. Building that ground truth is slow, unglamorous work that competes for budget with the next deployment, and it produces no demo.
The consequences differ by application. In back office automation the errors are often self-revealing, because an invoice coded to the wrong account eventually surfaces in a reconciliation. In customer-facing and advisory applications they are not, and a system that is subtly wrong in a consistent direction can operate for a long time before anything visible happens.
The gap also undermines the decisions built on top of it. A company that cannot measure its current system's performance cannot evaluate a replacement, which turns every model deprecation into an act of faith. It cannot make an honest build-versus-buy comparison. And it cannot answer the question its auditors and increasingly its regulators have started asking, which is not whether the company uses AI but whether it has evidence about how well it performs.
Some of the discipline is migrating from procurement inward. Buyers who negotiated standard contract terms discovered in the process that they needed acceptance criteria, and acceptance criteria require a test set. That is an unheroic path to good practice, and it is working better than exhortation.
Vendor benchmarks do not substitute, and treating them as though they do is the most common error. A published score describes performance on a public task set the vendor selected, which says close to nothing about performance on a particular company's documents, customers and edge cases.
The organizations that do this properly share an approach that sounds obvious and is rare: a fixed set of representative cases with known correct answers, scored the same way every time, run against every version and every candidate replacement, with results tracked over time. It costs real money to build once. It converts every subsequent question about the system from a debate into a measurement, which is the entire point, and it is the difference between deploying a capability and merely installing one.



