Fourteen gigawatts. That is the computing capacity the company says it will run by next year, up from roughly seven today. One gigawatt runs about 800,000 homes. It plans to spend up to $145 billion on AI infrastructure this year to get there.
The chip is the fourth generation of its MTIA accelerator program, designed with Broadcom and built by TSMC.
Its compute capacity doubles to 14 gigawatts by 2027, funded by up to $145 billion in 2026 infrastructure spend.
The news came from an internal memo reviewed by Reuters. In it, Meta admits something unusually candid for a company of its size: adopting the latest GPUs from Nvidia and AMD "has been a heavy lift, and it has cost us time."
Meta computing capacity
7 gigawatts in 2026, doubling to 14 by 2027 — about 11 million homes worth of power devoted to AI workloads. · Reuters, 2026
The chip that took five years to arrive
Meta's custom silicon program has a long and unglamorous history. The first MTIA chips were shown publicly in 2023, benchmarked against older Nvidia hardware rather than the newest parts on the market. Reuters described the broader effort as one that "has floundered since its launch more than half a decade ago."
Iris is meant to end that story. The chip cleared six weeks of bug testing without a single major architectural problem, fast enough that Reuters called it a notable milestone for a program that had struggled for years.
The chip is part of a four-generation roadmap the company unveiled in March 2026. The siblings — MTIA 300, 400, 450 and 500 — roll out on a roughly six-month cadence, which is double the industry's typical annual release cycle. MTIA 300 is already running its ranking and recommendation systems in production.
What Iris actually does
Iris is built to run the workloads that keep Facebook and Instagram alive: the ranking and recommendation systems that decide what appears in your feed, plus the generative AI features spreading across its apps. It handles training and inference for its Llama and Muse model families.
Nothing about this is meant to replace Nvidia. The memo is explicit that Iris works alongside the large volumes of Nvidia and AMD GPUs it keeps buying. It is a hybrid strategy: off-the-shelf GPUs for flexibility, custom silicon for cost efficiency at scale.
Why a hyperscaler builds its own chip at all
Confirmation criteria: Its cost per token on Llama workloads falls measurably once Iris capacity comes online in early 2027.
Broadcom's common thread
Meta did not design Iris alone. Broadcom handled the physical design and interface architecture, the same role it plays for Google's newest TPU and OpenAI's first custom chip. TSMC manufactures the silicon. For Broadcom, the arrangement is a reliable revenue stream: Mizuho analysts estimate the company collects $21 billion in AI-related revenue tied to Anthropic's use of Google TPU capacity in 2026, roughly doubling to $42 billion in 2027.
It formalized its Broadcom partnership through 2029, covering multiple MTIA generations. It also struck a multi-year deal with AMD for up to six gigawatts of Instinct GPUs, diversifying its compute supply away from a single vendor.
The supply chain underneath
Doubling computing capacity takes more than chip design. The memo shows it locked in long-term supply agreements with Samsung Electronics for memory chips, Sandisk for flash storage, and Sumitomo Electric for fiber-optic equipment. Memory and chip prices are climbing fast enough that the shortage story has moved upstream, from finished accelerators to raw memory dies.
This is the same bottleneck that reshaped the whole market in 2026. The scarce resource is no longer only model talent or user distribution. It is reliable access to power and chips at extreme scale.
| Parameter | Meta Iris | Nvidia GPU |
|---|---|---|
| Purpose | ✔ Internal inference for Meta apps | ✗ Merchant silicon, sold broadly |
| Workload fit | ✔ Tuned to Meta's ranking and Llama | ◐ General-purpose AI compute |
| Flexibility | ✗ Limited to Meta workloads | ✔ Very high, open market |
| Cost control | ✔ Cuts cost per token at scale | ◐ Dependent on vendor pricing |
Meta memo and analyst reporting, 2026
What this does to the market
Meta joins Google, Amazon and Microsoft in the custom-silicon club. Google's TPUs are on their seventh generation. Amazon deployed over 500,000 Trainium 2 chips. Microsoft runs Maia 200 internally.
The shift changes bargaining power. When a hyperscaler can threaten to run more workloads on custom silicon, Nvidia has less room to raise prices. That pressure is good news for everyone who pays for AI compute, from a large enterprise to a solo developer renting tokens.
Nvidia is not in immediate trouble. It still dominates the merchant accelerator market, with hyperscaler customers having achieved meaningful, not total, independence. But the trend line is clear: the biggest buyers of AI chips increasingly want to be their own suppliers.
Its cost per token on Llama workloads once Iris capacity ships in early 2027
Whether its six-month chip cadence holds through 2027 — a pace no hyperscaler has sustained
Nvidia's data center share as custom silicon scales across the big four cloud vendors
Memory supply: HBM and DRAM pricing as Meta, Google and Amazon all buy in volume
What happens to custom silicon by 2028?
Probability: 70% — the six-month release pace and Broadcom/TSMC relationships are already locked in; the main risk is execution slip, not intent.
✅ Arguments for
Meta already runs MTIA 300 in production; Iris is the next step, not a leap of faith.
Confirmation criteria: It reports lower per-token costs and keeps the roadmap on schedule.
❌ Arguments against
It does not own a fab — it depends on TSMC capacity like everyone else.
Disconfirmation criteria: a delayed generation, rising costs, or a retreat to merchant GPUs.
Development scenarios
🟢 Optimistic scenario (30%)
Implications: It undercuts rivals on AI serving costs; Nvidia's share erodes faster than expected.
🟡 Base-case scenario (50%)
Implications: steady efficiency gains; Nvidia keeps its lead but with thinner pricing power.
🔴 Pessimistic scenario (20%)
Implications: the capital sunk into custom silicon yields little; the cost curve turns against it.
As we wrote in August, inference silicon is becoming a one-trick wager: companies bet enormous sums that a single, narrow architecture wins the cost race. Meta's Iris is that bet applied at hyperscale, with the added twist that the company is both the chip's only customer and its primary user. That is the rare case where a custom accelerator gets to prove itself against a real, billion-user workload from day one.