The lesson enterprises took from the first wave of model deployments was concentration risk, and they acted on it. Contracts were rewritten to allow substitution, prompts were abstracted behind internal interfaces, and evaluation suites were built so a second provider could be qualified without a six-month project.
On paper, most large buyers can now move workloads between model vendors in weeks. The question nobody put on the diligence list is what those vendors have in common.
Substitution at the top, concentration underneath
Several of the providers a buyer would switch between run on the same handful of clouds. Many depend on the same accelerator supply, sometimes on the same generation of it. A growing share of enterprise traffic reaches models through a small set of inference platforms and gateways that sit between the application and whichever model is selected.
None of that appears in a contract that lists two suppliers. The switching option is real at the layer where it was tested and progressively less real at every layer below, and the layers below are where capacity actually binds.
Risk teams that have started mapping it describe an uncomfortable exercise. Asking a model vendor which cloud regions serve your traffic, which accelerator families they depend on and what their own concentration looks like produces answers ranging from detailed to politely refused, and the refusals cluster among the vendors with the most to disclose.
The failure mode is not a provider going out of business. It is correlated unavailability: a capacity crunch, a region incident or an allocation decision that affects several of your nominally independent suppliers at once, because they were never independent in the dimension that mattered. That is the same structural point index investors have been arguing about for years — diversification measured in names rather than exposures is not diversification.
Some of this is being addressed through the contract, in the terms buyers have been standardising across model agreements: disclosure of material subprocessors, notice periods on capacity changes, and the right to audit where inference physically runs. Those are useful and slow, and they lag the deployments they govern.
There is a second exposure inside the first. A model that is withdrawn takes its behaviour with it, and a buyer who has qualified an alternative on capability has usually not qualified it on output — the substitute answers differently, and every downstream threshold tuned against the original is now tuned against nothing. That is why deprecation has become a risk register item rather than a release note, and it compounds the concentration problem: the fewer genuinely independent options underneath, the fewer chances to absorb a withdrawal without rebuilding the evaluation work.
Spending patterns make the picture starker. As AI budgets moved out of experimentation and into core operating lines, the workloads stopped being ones a company could pause for a quarter while it re-qualified a supplier. Concentration matters more when the dependent process is invoicing than when it is a pilot.
The more practical mitigation is unglamorous. Keep one workload running continuously on the alternative rather than holding it in reserve, because a switching path that has never carried traffic is a document rather than a capability. Companies that run a genuine minority share on a second stack find out about the incompatibilities on a Tuesday instead of during an incident.
What makes this hard to prioritise is that the concentration has not cost anyone much yet. The outages have been short, the capacity crunches have resolved, and a risk that has not materialised competes badly for attention against a deployment that ships. That is usually the condition under which this class of exposure gets addressed late.
Topics aiprocurementriskcloud


