For two years the industry's scoreboard was model size, and bigger meant better meant news. The deployment data now tells a different story: the models actually shipping inside products and companies are getting smaller.
The logic is operational. A compact model tuned to a narrow job, routing support tickets, extracting fields from freight documents, summarizing case notes, matches or beats a giant generalist on that job while costing a fraction to run. It can live on a company's own hardware, which resolves at a stroke the privacy and data-residency questions that stall enterprise deals.
The portfolio approach
The emerging corporate pattern is a portfolio: a frontier model rented for the hardest reasoning, a mid-sized workhorse for general tasks, and a fleet of small specialists embedded in workflows. Costs concentrate where capability is actually needed instead of being smeared across every API call.
Hardware trends reinforce the shift, as laptops and phones ship with accelerators that make local inference ordinary, and spending on AI infrastructure increasingly splits between training giants and serving dwarfs.
None of this diminishes the frontier, which still defines what is possible and generates the distilled knowledge smaller models inherit. It does redistribute the money. The frontier is a research budget. The small model is a cost of goods sold, and the second category is always, eventually, larger.
The cost argument has become the main one rather than a secondary consideration, as inference moves from footnote to line item in budgets that were built when usage was a pilot.
Related reporting has traced AI Agents Take Over the Back Office, Quietly and States Are Writing the AI Rulebook, One Legislature at a Time.



