The memory wall is where AI inference is now decided. Positron AI closed an $875 million Series C on September 10, 2026 — at a $5 billion post-money valuation, up from $1 billion seven months earlier — on a contrarian claim: that the cheapest way to serve large models is to stop paying for high-bandwidth memory.

Every dollar of inference cost now runs through memory bandwidth. Training is largely solved at the frontier. Serving a model to millions of users is not. The binding constraint is not raw compute but how fast weights and context move to that compute, and at what power cost.

Positron's answer skips HBM entirely. Its second-generation chip, Asimov, pairs a systolic-array datapath with 288 GB to 2,304 GB of LPDDR5X per die — phone-grade memory, placed next to the compute, instead of the expensive stacked memory that Nvidia's accelerators depend on.


The memory wall moved down the stack

For three years the industry measured AI progress in FLOPs. That metric is now the wrong one to watch. Inference — running a finished model, not training a new one — is the majority of new compute being bought, and it is dominated by the cost of moving data. A GPU that can multiply faster spends most of its time idle while it waits for weights and key-value cache to arrive.

HBM solved part of that problem by stacking memory vertically next to the die. The tradeoff is price and supply. High-bandwidth memory is expensive, capacity-constrained, and increasingly gated by advanced packaging processes — CoWoS and HBM4 allocation are now the scarce inputs for the entire accelerator market.

Positron's bet is that for the specific job of serving transformers, you can trade peak bandwidth for capacity and cost. Put a large pool of cheap LPDDR5X next to a custom datapath, keep the model weights resident in that co-located memory, and serve very long contexts from a single node. As we wrote in October, the memory wall is the ground where inference startups now fight; Volantis raised $88M to attack it with photonics. Positron attacks it with commodity memory instead.

$875M Series C raise

Positron AI's Series C at a $5B valuation

Closed in two tranches: $375M at a $3.5B pre-money valuation, plus up to $500M in a C-1. · Company release, Sep 2026

Asimov bets against HBM

Asimov is a systolic array: a grid of identical compute modules, each with its own co-located memory. The design keeps weights close to the arithmetic units that consume them, which cuts the round trips that dominate HBM-centric serving. It tapes out on TSMC's N3P node at the end of 2026, sixteen months after design work began, with production targeted for the second half of 2027. Each die carries between 288 GB and 2,304 GB of LPDDR5X memory.

Titan, the system built around it, links four to eight Asimov chips into a single node. Positron says a rack will serve models beyond 16 trillion parameters with context windows past 10 million tokens — workloads that today require spreading a single model across many expensive GPU nodes.

MetricPositron Asimov / TitanNvidia GB300 NVL72
Memory type ✔ LPDDR5X, co-located ◐ Stacked HBM3e
Memory per chip ✔ 288–2,304 GB ✗ Fixed, HBM-bound
Status ◐ Pre-production (2027) ✔ Shipping at scale
Cost claims ◐ Simulation-only ✔ Deployed benchmarks

Positron company claims vs. Nvidia shipping hardware. Cost claims are simulation-derived, not measured output.

What the money actually buys

The round was split, and the split matters. A $375 million Series C at a $3.5 billion pre-money valuation was co-led by NEA, Andra Capital, Atreides Management, Valor Equity Partners, and Dylan Patel's SemiAnalysis Capital. A C-1 tranche of up to $500 million was anchored by NEA and Jim Clark, the co-founder of Silicon Graphics and Netscape. The money funds three specific things: the Asimov tapeout, a 2 MW-plus engineering data center and emulation platform, and the production ramp of Titan.

Positron is not selling simulations. Atlas, its first-generation system, is already deployed at Oracle Cloud Infrastructure, with more than 50 racks in production. Parasail, a partner, runs its own inference service on that capacity. Additional Atlas customers include Jump Trading and i3d.net.

Our focus now is to tape out Asimov, bring Titan to production, and scale manufacturing to meet the demand in front of us.— Mitesh Agrawal, CEO, Positron AI

That is the part of the pitch that has shipped. The part that has not is the headline efficiency number. Positron says simulations show a rack of its silicon processing up to 26 times as many tokens per dollar as Nvidia's GB300 NVL72, and roughly 5 times as many tokens per watt as Nvidia's Rubin. Both figures come from modeling, not from a production deployment. Treat them as a design target, not a benchmark.

🎯
Inference cost is now a memory problem, not a compute problem.

Positron is betting that cheap, co-located LPDDR5X beats expensive HBM at serving long-context models.

The valuation is a 2027 wager: the product that justifies it is not in production yet.

The risk inside the pitch

The memory-first argument has a real hole in it. Skipping HBM caps peak bandwidth. For workloads that are bandwidth-bound rather than capacity-bound — high-throughput, short-context, batch serving — that tradeoff is a disadvantage, not an edge. Positron's answer is that long-context agentic work is the growth segment, and there its capacity advantage compounds.

The competitive field is also crowding fast. Nvidia is not standing still on inference efficiency. AMD is integrating hardwired inference silicon into its accelerator roadmap. And a generation of inference startups — Volantis, SiMa.ai, Rebellions, Euclyd — is attacking the same cost curve from different directions. Positron's differentiator is the memory choice, and it is a bet that only pays if the market's workloads shift the way it expects.

The tapeout is the next hard checkpoint. A design that works in emulation and fails in silicon would reset the story. Positron has to convert $875 million and a $5 billion valuation into a working chip on a 2027 timeline, against incumbents with shipping products and endless memory budgets.

Does memory-first inference win the serving layer by 2027?

🔮
By the end of 2027, at least one memory-first inference system will hold a measurable share of production long-context serving.

Probability: 55% — the demand curve for long-context agentic workloads is real, but memory-first has to survive a first tapeout and prove its numbers outside simulation.

✅ Arguments for

HBM supply is the bottleneck on every accelerator roadmap; LPDDR5X sidesteps it.

Long-context agents reward capacity over peak bandwidth.

Confirmation criteria: Asimov tapes out on schedule and a named customer runs production inference on Titan in 2027.

❌ Arguments against

Lower bandwidth loses on short-context, high-throughput serving.

The 26x claim is simulation-derived, not shipped.

Falsification criteria: Asimov misses its tapeout window, or deployed Titan fails to beat HBM systems on tokens per dollar.

Development scenarios

🟢 Optimistic scenario (30%)

Asimov tapes out on time and Titan lands real production customers in 2027.

Consequences: Positron becomes the default alternative to HBM-centric serving and the $5B valuation looks early.

🟡 Base scenario (50%)

The chip ships in 2027 but wins a niche: long-context and research serving, not the mainstream.

Consequences: A durable business, smaller than the round implies.

🔴 Pessimistic scenario (20%)

Silicon slippage or a bandwidth penalty that the long-context market never fully offsets.

Consequences: Positron is re-priced toward its Atlas business, and the memory-first thesis recedes.
Chipmaker Positron nabs $875M to speed up inference with consumer-grade memory
Technical breakdown of the Asimov systolic array and the decision to drop high-bandwidth memory for co-located LPDDR5X.
The clearest account of what the chip architecture actually changes.
Positron AI raises $875m at a $5bn valuation to develop next-generation hardware
Tapeout timing on TSMC N3P, the two-tranche structure of the round, and the 2027 production plan.
Datacenter infrastructure angle: where the 2 MW engineering build fits.
Positron AI raises $875 million at a $5 billion valuation
Company statement on the Atlas deployment at Oracle Cloud and the financing's use of proceeds.
Primary source text for the round, the valuation, and the customer list.