Enterprise token costs fell 67% in a year. Inference spending is still forecast to triple, from $120B in 2025 to $885B in 2030. Both figures describe the same market, and the distance between them explains why four different ways of running inference now compete for the same workload.
The fourth of those arrived on 2 September 2026, when Equinix announced Inference Exchange at its first Equinix Horizon customer event: NVIDIA's validated enterprise reference architecture, Together AI's serving platform with more than 200 open-source models, and Equinix's own data centers and fabric. General availability is set for the first quarter of 2027.
WHERE ENTERPRISE INFERENCE RUNS
──────────────────────────────────────────────────────────────────
Jun 2020 Aug 2026 Sep 2026 Sep 2026 Q1 2027
the baseline $240M Inference 67% cost Equinix and
for a IBM + Exchange drop, 59% IBM clusters
600x drop Together AI announced off-cloud go live
──────────────────────────────────────────────────────────────────
Chronology compiled from the Equinix and IBM announcements of 11 August and 2 September 2026, plus arXiv 2603.28576 and Futurum Group tracking data, September 2026.
Two curves pulling the same bill
Serving costs are collapsing while usage climbs. A 2026 arXiv study of inference token pricing measures roughly a 600-fold decline since June 2020, with economy-tier serving landing near $0.10 per million tokens. That deflation is what funds the volume growth behind the $885B figure.
Drop in token price
Economy-tier inference serving now sits near $0.10 per million tokens, and enterprise cost per workload fell 67% year over year. · arXiv 2603.28576; Futurum Group, 2026
Demand is the second curve. The same arXiv work finds reasoning-heavy traffic priced at a 31.5× premium over the economy tier on average, because a thinking model burns far more tokens per answer. Futurum Group's September 2026 enterprise tracking puts reasoning workloads up 219%, and agentic workloads can consume up to 100 times the tokens of a single chat turn.
Price gap over economy tier
Reasoning traffic costs far more per answer because it burns more tokens, and that gap is what specialists sell against. · arXiv 2603.28576; Futurum Group, 2026
For an investor the arithmetic is unfriendly to anyone selling raw tokens and friendly to whoever owns the scarce inputs around them: power, cooling, sites with grid access, and low-latency fabric. Compute is turning into the cheaper half of the bill.
Model one: the hyperscaler default
AWS, Azure and Google Cloud still win the default enterprise workload, and the reason is proximity to everything already bought. Identity, data pipelines, vector stores, audit logs and the security team all sit inside one account, and inference is a single API call away. Capacity is sold through reservations and committed-use contracts, which turns a variable cost into a fixed line a finance team can plan against.
The trade is lock-in. The more of the stack a company moves onto rented GPUs, the harder the exit becomes, and the exit is what gets tested when a token price moves. Hyperscalers also place new accelerator generations first, which matters for reasoning models with specific memory requirements.
Model two: the specialist neocloud
Specialists sell open-model serving at the deflationary end of the market and enterprise bare-metal clusters at the top. Together AI is the clearest case. On 11 August 2026 IBM signed a multi-year agreement worth $240M to stand up a dedicated cluster of NVIDIA HGX B300 systems with Spectrum-X networking on IBM Cloud, available in the first quarter of 2027.
Read one way, an enterprise bought reserved capacity from a hyperscaler. Read another, a hyperscaler bought a specialist's software stack and model catalogue to defend the workload. Both readings are consistent with the announcement, and the pricing power sits in whoever owns the catalogue.
The economics live in the middle. A specialist can serve an open model below a hyperscaler's list price because it runs fewer clouds and fewer compliance regimes around the same GPU. What it cannot easily offer is the identity and audit layer, and for a regulated buyer that often decides the contract.
Model three: the neutral exchange
The third model has no obvious ancestor. Inference Exchange runs NVIDIA's validated enterprise reference architecture across more than 280 Equinix data centers in 77 metros, reachable to clouds, networks and AI providers through Equinix Fabric. The program was shown in preview at Equinix Horizon alongside Fabric One, with availability scheduled for the first quarter of 2027.
The pitch is neutrality. A landlord that already owns power, cooling and cross-cloud connectivity can sell inference the way it sells colocation: by placing the workload close to the data and the users while staying out of the model business. Equinix is the one party in that arrangement no hyperscaler region can replace, and it is the asset being sold.
The infrastructure chosen now will determine a company's competitiveness for years to come.— David Fox-Martin, CEO, Equinix
Two details shape the margin. The exchange supports both multitenant deployments for shared efficiency and dedicated single-tenant environments, so one footprint can serve a startup paying by the token and a bank buying a private cluster. And it monetises an externality: cheap tokens push value upward into the site, the power contract and the network path.
Token deflation turns compute into a commodity input and leaves the scarce assets where they already sit: contracted power, liquid cooling, and a cross-cloud network path. Futurum Group expects 59% of AI workloads to run outside public hyperscaler cloud, which sizes the wedge this model is aimed at.
The obvious risk is that hyperscalers answer with the same neutrality, having sold connectivity to their own regions for two decades. The counter is that Equinix already sits between those clouds as a landlord paid by both of them.
Model four: on-premises and the edge
The fourth model is the oldest and the least fashionable, and it survives where data cannot move: sovereign deployments, regulated industries, and sites that must keep inference close to a factory floor or a clinical system. Engineering.com's reporting on the Equinix launch framed sovereign AI as a named use case for exactly this reason.
The cost is utilisation. A dedicated cluster idles when demand is seasonal, and the operator carries depreciation, the power contract and the refresh cycle without any cloud provider's elasticity. That is why sovereign deployments tend to be framed as insurance rather than as optimisation.
| Delivery model | Commitment | Governance | Where it wins |
|---|---|---|---|
| Hyperscaler cloud | ◐ Reserved or committed use | ✔ Full identity and audit stack | Workloads already inside one account |
| Specialist neocloud | ✔ Pay per token or bare metal | ✗ Thin identity layer | Open models at the deflationary end |
| Neutral exchange | ◐ Multitenant or single tenant | ◐ Operator-managed, cloud-neutral link | Data and users spread across metros |
| On-premises and edge | ✗ Full capex and refresh cycle | ✔ Data never leaves the site | Regulated and sovereign workloads |
Editorial comparison of commitment, governance and fit. Commitment and availability drawn from the Equinix and IBM announcements of 2 September and 11 August 2026; governance assessed against each model's standard enterprise contract, September 2026.
The arithmetic that decides it
Unit economics now favour the buyer. With token prices down roughly 600-fold and enterprise cost per workload down 67% year over year, the deciding variable is the fixed cost around the GPU: power contracted years ahead, cooling, and the distance between the workload and the user. That is why the two deals announced within three weeks of each other, IBM's $240M with Together AI on 11 August and Equinix's exchange on 2 September, are both about sites and networks.
Demand elasticity pulls the other way. If reasoning workloads grow 219% and agentic traffic burns up to 100 times the tokens, the total bill keeps climbing while the unit price falls. Falling unit prices with a rising total bill describe a market where margin moves upward, toward whoever controls the scarce input.
Turning points we flagged earlier
Two bets in this section pointed at the same shift before the announcements did. Rebellions' sovereign inference strategy, covered on 25 September, is the on-prem and regulated model with a regional chip supply chain attached. Runware's 1 MW inference pod, covered on 22 September, attacked the same problem from the opposite end: bring the data centre to the customer in a shipping container, with facilities costs running far below a gigawatt-scale build.
Read together, both are attempts to break the rule that inference has to happen in someone else's building. The Equinix announcement is the landlord's answer, and it lands in the same quarter as both.
Q1 2027 for both Equinix Inference Exchange and the IBM cluster with Together AI, then whether any hyperscaler answers with cloud-neutral fabric of its own, then whether agentic token growth outruns the 67% annual price decline.