Two-thirds of AI compute now runs inference rather than training. The number has flipped in eighteen months, and with it the entire economics of the industry. The companies that print the tokens are no longer the ones that matter most to an investor's returns.
Value is accruing at the inference and infrastructure layer, not uniformly across the stack.
The neocloud model wins the inference tier on price, but carries debt and depreciation risk the hyperscalers do not.
The shift is not a chart curiosity. It decides which balance sheets you should be reading, and which "AI infrastructure" stories are actually about returns.
The flip that rewrote the stack
For most of the last three years the infrastructure conversation was a training story. How many accelerators can a cluster hold. How to keep a thousand GPUs fed with data without stalling. That problem is largely solved for the frontier labs, and it was never the constraint for the other 99 percent of teams.
In 2026 the daily fight is inference. Latency per token. Throughput under concurrent load. Cost per request at the tail of the traffic distribution. Your model is already fine-tuned. Now it has to answer ten thousand requests an hour with sub-500-millisecond response times, and the bill has to scale with revenue rather than ahead of it.
This matters for capital allocation because training runs once per model version. Inference runs on every single user request. The workload that used to be the warm-up is now the plant. As we wrote in July on the hyperscaler capex guide, the buildout was never really a chip story. It is a recurring-spend story, and inference is where that spend now lives.
The consequence is a stack that looks less like a ladder and more like a set of distinct businesses, each with its own margin profile and its own winner.
Seven layers, one bill
A production AI stack in 2026 spans roughly seven independent markets. Picking the right vendor at each layer, and knowing where the layers touch, is what separates teams that ship on budget from teams that burn compute on mismatched tooling.
At the bottom sits compute: the GPUs, the custom accelerators, the bare metal. Above it, the data and vector layers that hold the context an agent retrieves. Then orchestration, the serving and inference layer that turns a model into an API, the fine-tuning and training control plane, observability and evaluation, and finally governance and guardrails.
Each layer has its own economics. Compute is a commodity with a financing tail. Serving is a margin game won on utilization. Observability is a recurring subscription. Governance is a tax you pay to avoid a headline.
The practical point for anyone evaluating this sector: do not underwrite "AI infrastructure" as one thing. The returns live in specific layers, and they diverge sharply.
Where the value accrues
The cleanest framing comes from the analysts who model total cost of ownership (TCO) for a living. Across 2023 to 2025, almost all the value in AI was captured by the infrastructure layer. In 2026 that is narrowing further, and the split is sharp.
At the model-lab tier, the numbers are striking. Anthropic's annualized revenue moved from roughly $9 billion to more than $44 billion in a single year, while gross margins on its inference infrastructure climbed from 38 percent to over 70 percent. The labs are capturing value they held almost none of eighteen months ago.
Below them, the silicon and foundry layer vents value into every vertical. The accelerator makers and the leading foundry sit on pricing power that training alone never justified. Inference at scale only deepened it.
The contested middle is the neocloud and inference-serving tier. This is where the most capital is being raised and the least is yet proven. It is also where a private investor can still take a position before the margins are settled.
Foundry and accelerator supply: pricing power, but capital-heavy and cyclical.
Model labs: capturing the largest margin expansion of any tier.
Neoclouds and inference serving: highest capital inflow, unproven free cash flow.
Orchestration and observability: steady recurring revenue, lower ceiling.
The neocloud wedge
The neocloud is a specialized GPU cloud built for AI workloads rather than general compute. Where a hyperscaler sells you a virtual machine and a billing surprise, a neocloud sells you a GPU hour with networking and storage wrapped in, at a single transparent rate.
For years this was a marginal category. In 2026 it is the fastest-moving part of the buildout. Groq, which began as a chip company, raised $350 million in August to pivot fully into inference cloud, targeting a jump from 54 megawatts of capacity to more than 200 megawatts in 2027 and serving over six million developers. CoreWeave reported active capacity above 1.5 gigawatts with roughly 4.2 gigawatts contracted, and guided annual capital spend between $35 billion and $39 billion.
The wedge is price. Reserved next-generation accelerator capacity holds a meaningful discount to the hyperscaler equivalent, and the inference tier is contested for the first time since the buildout began. Specialty silicon from multiple vendors now lands in production inference, not just training.
NVIDIA, which supplies most of these clouds, has also changed the financing model. Through a partner program it backs multi-tenant "AI factories" with revenue-sharing and credit support, letting operators bring capacity online without sitting through years of site selection, power procurement, and construction. Firmus is building a campus expected to scale toward 360 megawatts and up to 170,000 GPUs. Sharon AI is deploying tens of thousands of the latest accelerators for sovereign compute.
The strategic question is whether neoclouds are a durable business or a leveraged bet on hardware that depreciates faster than the debt underwriting it.
For a private investor, the inference tier is the most interesting because it is the least settled. The foundry and accelerator layer is effectively a duopoly with pricing power but limited entry points for non-strategic capital. The model labs are mostly already valued as category winners. The serving and neocloud tier is still a fight, which means it is still a place where selective capital can buy a position at a reasonable entry rather than a consensus multiple. The same contestability that makes the margins uncertain is what makes the upside open.
What the buildout actually costs
The reason this sector attracts private capital is simple. It is the most capital-intensive category in technology. The financing structures are as important as the technology.
Capacity is funded through a mix of equity, GPU-backed debt, and asset-backed securities collateralized on the chips themselves. Underwriters increasingly require contracted revenue visibility, often a multi-year take-or-pay commitment from a hyperscaler or frontier lab, before they release capital. Build speculatively and the capacity strands the moment a GPU allocation slips.
Then there is power. The interconnection queues at major grid operators run years long. A new data-center site is frequently a 2029 or 2030 project unless it can redirect existing generation. The binding constraint on the entire buildout is no longer silicon. It is electrons and permits.
For an investor, that reframes the diligence. The interesting infrastructure trend of 2026 is not the megasite. It is the on-prem and colocation rebuild, where regulated buyers with data they cannot move, latency they cannot meet across a public-cloud network, and audit demands they cannot satisfy as a tenant, buy their own accelerated clusters. The hyperscalers built dedicated-region products to meet some of this. The on-prem AI factory pattern meets the rest.
The case against the hype
The bull case writes itself. The bear case is the one that protects capital.
First, hardware depreciation. Accelerators are financial liabilities the moment they are racked, and the product cycle shortens every year. A fleet that looked flagship in January can look mid-tier by the following spring. Operators that finance capacity on debt are exposed to exactly this gap.
Second, utilization. Inference economics are won on GPU utilization, not on list price. A reserved cluster that idles at 40 percent utilization quietly costs far more than its headline rate suggests. The operators that win are the ones treating infrastructure planning as a multi-year capital project rather than a procurement line item.
Third, concentration. A neocloud's survival often depends on one or two anchor tenants and on continued supply from a single accelerator vendor. That is a fine structure while demand outruns supply. It is a fragile one the moment supply normalizes or an anchor tenant builds its own capacity.
None of this makes the sector uninvestable. It makes the selection bar high. The returns will not be evenly distributed.
Serving is where the margin is made
Of all seven layers, the inference-serving layer is the one most investors misunderstand and the one most likely to decide returns. Training optimizes for raw throughput and checkpoint reliability. Serving optimizes for cold-start latency, concurrent request throughput, and per-request cost. None of those share an optimization target, yet teams keep building serving infrastructure as if it were a training problem.
The framework choice sets the cost floor. Modern continuous-batching servers deliver roughly 23 times the throughput of naive serving on the same hardware. That single architectural decision is the difference between a unit cost that compounds in your favor and one that quietly eats the margin. Below the framework sits routing: matching workload class to capacity class. Long-running training jobs belong on reserved capacity. Batch and exploratory work belongs on spot. Latency-sensitive inference belongs on dedicated hardware. The arbitrage is in the matching, not in the hardware.
A newer wrinkle is disaggregated inference, where the prefill and decode phases of a request are split across different accelerators. It sounds like an internal optimization detail. It is actually a margin lever, because the two phases have opposite resource profiles and forcing them onto one machine wastes both. Operators that get this right can price a token cheaper than competitors running monolithic serving, and the gap shows up directly in utilization.
This is also where the open-weight model ecosystem changes the calculus. Managed inference APIs abstract the GPU away entirely: you call an endpoint, pay per token, and never think about utilization. Self-hosted serving runs on your own cloud and costs only compute, which is materially cheaper at scale but demands engineering discipline the API hides from you. The break-even point has moved. A year ago, self-hosting only made sense past a certain volume. In 2026, with inference pricing under structural pressure from falling commodity rates, the self-host line is crossing for far smaller workloads.
The investment read is that serving is a margin game won on utilization and routing, not on owning the most GPUs. The operators worth underwriting publish their utilization, not just their announced capacity. Capacity is a promise. Utilization is the business.
How to read the financials
Most of the companies in this sector report soaring revenue and soaring losses in the same breath. The line that matters is not either. It is the relationship between announced capacity, contracted backlog, and free cash flow.
Start with contracted backlog. A multi-year take-or-pay commitment from a hyperscaler or frontier lab is what underwrites the GPU-backed debt that funds the buildout. Without it, capacity is speculative and the financing door closes. Treat backlog as the real revenue, and announced capacity as the ambition.
Then watch capital intensity against free cash flow conversion. The sector's capital spend is measured in tens of billions for the leaders, and depreciation on accelerators is front-loaded. A company guiding $35 billion to $39 billion of annual capex is making a bet that utilization and pricing hold long enough to amortize it. If either slips, free cash flow stays negative for years and the equity finances the gap.
GPU-backed asset-backed securities are the canary. When underwriters tighten the contracted-revenue visibility they require, or when spreads on this paper widen, it signals that the debt market doubts the cash flows underneath the chips. That is an earlier warning than any earnings miss.
Finally, separate the infrastructure owners from the infrastructure users. An operator that owns the GPUs and the power carries the depreciation and the debt. A software layer that orchestrates or observes that capacity carries far less capital risk and often better margins. The same "AI infrastructure" label hides both, and the balance sheet is the only honest decoder.
What to track from here
Utilization rates at the large neoclouds, not just announced capacity.
The share of AI compute spend going to inference versus training quarter over quarter.
First signs of accelerator oversupply or a softening in GPU-backed debt markets.
Whether on-prem AI factory deployments expand beyond regulated industries.
The training-to-inference flip is not a forecast. It is the current state of the plant, and it is what should anchor any position in this sector. Read the stack layer by layer, follow the recurring spend, and treat announced capacity as a promise until utilization proves it.
The opportunity is real and the capital inflow is historic. So is the chance of mistaking a leveraged hardware bet for a software-margin business. The investors who do well here will be the ones who can tell the difference from a balance sheet, not from a press release. The buildout will continue for the rest of the decade. The returns will not be shared equally, and the stack is where that divergence is decided.