50%. That is the cost reduction OpenAI's new inference chip is built to deliver, token for token, against the general-purpose GPU stack the company currently rents. The chip, codenamed Jalapeño, is an application-specific integrated circuit (ASIC) co-developed with Broadcom, joining a club that already holds Google, Amazon, Microsoft and Meta. Every member built its own silicon for the same reason: cost-per-token is now the defining metric in AI infrastructure.
Five hyperscalers now run their own inference silicon, and each one is competing with NVIDIA on the one number that matters.
The battle has moved off peak flops and onto dollars per token, which shifts the economics of every AI product with a usage curve.
The launch is easy to read as one company fixing its own cost problem. It is bigger than that. OpenAI is the largest single consumer of AI compute in the world, and its decision to build a bespoke part signals where the market is headed. When the biggest buyer stops treating GPUs as the only option, the GPU premium starts to erode.
Jalapeño inference cost target
The custom ASIC, co-developed with Broadcom, is built for LLM inference rather than general compute, with a stated goal of roughly halving inference cost. · OpenAI, Broadcom, 2026
The custom-silicon club
Google's TPU line, Amazon's Inferentia and Trainium, Microsoft's Maia, Meta's MTIA, and now OpenAI's Jalapeño. Five in-house inference chips competing with the merchant GPU. · Silicon Analysts, 2026
What is growing: the custom-silicon club
Custom inference chips moved from experiment to strategy in about three years. Google started the pattern with the Tensor Processing Unit (TPU), a part designed for exactly one job and priced for exactly one metric. Amazon followed with Inferentia for inference and Trainium for training. Microsoft built Maia for Azure. Meta's in-house training and inference accelerator (MTIA) now runs part of the recommendation workload that once sat on merchant GPUs. Jalapeño is the fifth entry, and it is the loudest signal yet because it comes from the company with the largest compute bill in the industry.
The logic is consistent across all five. A GPU is a general-purpose part. It trains models, it serves them, it renders, it accelerates a dozen other workloads. That breadth carries a tax. The silicon spends transistors on generality nobody needs at the moment of inference. A custom ASIC strips that tax away. It implements only the operations an LLM actually performs at serving time, which frees die area, memory bandwidth and power for the work that matters. The result is a lower cost per token on the specific workloads the owner cares about.
This is the same logic every big compute buyer discovered in the 2010s. Search engines moved search ranking onto field-programmable gate arrays (FPGAs) and custom parts. Cryptocurrency miners moved proof-of-work onto ASICs and left GPUs for gamers. Networking vendors replaced merchant switch chips with custom silicon. The pattern is not new. What is new is that the largest AI company in the world has joined it.
What is fading: the GPU premium
NVIDIA built a fortress on the insight that a single powerful chip could serve every AI workload. That thesis is still true for training. Foundation models are built with GPUs, and nobody in the club above is seriously challenging that yet. Inference is where the fortress walls are thinner.
The economics are easy to state. When a model is trained once and served a billion times, the training cost gets amortised into insignificance. The serving cost becomes the entire bill. Every percentage point of inference cost reduction flows straight to margin, and for OpenAI, with inference traffic at hyperscale volumes, half of that bill is a number with serious zeros behind it.
The GPU premium is fading for a second reason. Pricing power depends on scarcity, and custom silicon attacks scarcity from two directions at once. It reduces the demand for merchant parts, and it gives the buyer a credible outside option in every negotiation. A hyperscaler that can walk away from a GPU generation is a very different customer from one that cannot.
What is new: Jalapeño in detail
Jalapeño is an inference-focused design, and the division of labour between OpenAI and Broadcom is itself a new development. Broadcom has spent years building custom networking and accelerator ASICs for hyperscale customers, most visibly for Google's TPU program. OpenAI brings the model requirements, the workload profiles and the deployment pipeline. Broadcom brings the chip design, the packaging and the manufacturing relationship. The partnership model means it does not have to stand up a silicon team from scratch to own its inference stack.
The target is explicit: roughly half the cost per token. That kind of number does not come from incremental tuning. It comes from a redesign around inference-specific arithmetic, memory layout and serving patterns, with every transistor justified against the serving-time workload rather than the training run. It is the difference between a tool built for the job and a tool borrowed from another job.
What remains open is the deployment schedule. Custom silicon is a multi-year program, and the gap between unveiling and meaningful volume is where such projects usually stumble. The same discipline that makes ASICs cheap per token also makes them unforgiving: a custom part is wrong the moment the workload changes in a way the designers did not anticipate. The club's members have accepted that risk because the prize is control of their own cost curve.
What the shift means for who pays
Follow the cost curve to its end and the conclusion is uncomfortable for the merchant-GPU business model. The hyperscalers who buy the most AI compute now all hold a second option. Even if a custom ASIC never reaches full volume, the credible threat of one changes the negotiation. A customer who can walk away from the latest GPU generation is no longer captive to its price.
For the companies building on top of that infrastructure, the direction of travel matters more than the timing. Inference costs have been falling every year, and the arrival of custom silicon accelerates the slope. Every AI application with a usage curve inherits that decline. Products that are uneconomic at today's prices stop being uneconomic at half the cost per token. That is where new business models appear, and it is where a patient investor should look.
The winners are not obvious from the outside. Chip suppliers with exposure to the custom-accelerator design cycle gain a revenue stream that is less cyclical than merchant GPU sales. Cloud platforms that own their silicon capture the margin directly. Software companies that price per token or per request ride the cost decline automatically, even though the decline itself is not their doing. The losers are the pure renters who neither own silicon nor pass falling costs to customers.
None of this is guaranteed. Custom silicon has failed before, and the difference between an unveiled chip and a shipped one is measured in years and billions. But the strategic direction is set. The largest buyer in the industry has decided that owning the silicon is the way to control its own margin, and the rest of the market will price around that decision. The interesting question is no longer whether custom inference silicon works. It is which company turns the advantage into a durable one.
ASIC versus GPU: the cost-per-token math
| Parameter | Custom inference ASIC | General-purpose GPU |
|---|---|---|
| Design goal | ✔ Minimum cost per token on LLM serving | ◐ Broad compute, training plus inference |
| Cost per token | ✔ Target roughly 50% lower | ✗ The incumbent baseline |
| Flexibility | ✗ Narrow, workload-specific | ✔ Adapts to new model architectures |
| Time to market | ✗ Multi-year co-design cycle | ✔ Off the shelf, yearly cadence |
| Supply control | ✔ Owned roadmap and allocation | ✗ Depends on vendor roadmap |
The table shows why the war is not a rout. Custom silicon wins the metric it was built for, and loses everything else. A GPU can pick up a new model architecture without a design cycle. A custom ASIC cannot. The merchant part retains optionality, and optionality has value when the underlying technology is still moving fast. That is the honest summary of the trade: efficiency now versus flexibility later.
Investors should read the balance the same way. The companies that control their own cost curves gain a durable margin advantage in serving. The companies that rent GPUs at merchant prices remain exposed to the pricing power of a single supplier. Neither side is wrong. They are simply making different bets on how fast the underlying technology changes.
When it moves a meaningful share of its inference off rented GPUs onto Jalapeño, the cost-per-token claims become auditable in real output prices.
Broadcom's custom-accelerator revenue now has a second anchor customer alongside Google's TPU program.
NVIDIA's next architecture reveals whether the response is an inference-specific stock keeping unit (SKU) or a continued general-purpose push.
Each of the other four club members either expands its ASIC scope or quietly stays on merchant parts.
The quiet conclusion is the durable one. AI infrastructure is no longer a one-vendor story. The largest buyer in the industry has decided that owning silicon is cheaper than renting it, and that decision reshapes the cost structure of the entire market. The price of intelligence is falling, and it is falling because the people who pay the biggest bills decided to stop waiting for someone else to lower it.