The memory wall is where AI inference is now decided. Positron AI closed an $875 million Series C on September 10, 2026 — at a $5 billion post-money valuation, up from $1 billion seven months earlier — on a contrarian claim: that the cheapest way to serve large models is to stop paying for high-bandwidth memory.
Every dollar of inference cost now runs through memory bandwidth. Training is largely solved at the frontier. Serving a model to millions of users is not. The binding constraint is not raw compute but how fast weights and context move to that compute, and at what power cost.
Positron's answer skips HBM entirely. Its second-generation chip, Asimov, pairs a systolic-array datapath with 288 GB to 2,304 GB of LPDDR5X per die — phone-grade memory, placed next to the compute, instead of the expensive stacked memory that Nvidia's accelerators depend on.
The memory wall moved down the stack
For three years the industry measured AI progress in FLOPs. That metric is now the wrong one to watch. Inference — running a finished model, not training a new one — is the majority of new compute being bought, and it is dominated by the cost of moving data. A GPU that can multiply faster spends most of its time idle while it waits for weights and key-value cache to arrive.
HBM solved part of that problem by stacking memory vertically next to the die. The tradeoff is price and supply. High-bandwidth memory is expensive, capacity-constrained, and increasingly gated by advanced packaging processes — CoWoS and HBM4 allocation are now the scarce inputs for the entire accelerator market.
Positron's bet is that for the specific job of serving transformers, you can trade peak bandwidth for capacity and cost. Put a large pool of cheap LPDDR5X next to a custom datapath, keep the model weights resident in that co-located memory, and serve very long contexts from a single node. As we wrote in October, the memory wall is the ground where inference startups now fight; Volantis raised $88M to attack it with photonics. Positron attacks it with commodity memory instead.
Positron AI's Series C at a $5B valuation
Closed in two tranches: $375M at a $3.5B pre-money valuation, plus up to $500M in a C-1. · Company release, Sep 2026
Asimov bets against HBM
Asimov is a systolic array: a grid of identical compute modules, each with its own co-located memory. The design keeps weights close to the arithmetic units that consume them, which cuts the round trips that dominate HBM-centric serving. It tapes out on TSMC's N3P node at the end of 2026, sixteen months after design work began, with production targeted for the second half of 2027. Each die carries between 288 GB and 2,304 GB of LPDDR5X memory.
Titan, the system built around it, links four to eight Asimov chips into a single node. Positron says a rack will serve models beyond 16 trillion parameters with context windows past 10 million tokens — workloads that today require spreading a single model across many expensive GPU nodes.
| Metric | Positron Asimov / Titan | Nvidia GB300 NVL72 |
|---|---|---|
| Memory type | ✔ LPDDR5X, co-located | ◐ Stacked HBM3e |
| Memory per chip | ✔ 288–2,304 GB | ✗ Fixed, HBM-bound |
| Status | ◐ Pre-production (2027) | ✔ Shipping at scale |
| Cost claims | ◐ Simulation-only | ✔ Deployed benchmarks |
Positron company claims vs. Nvidia shipping hardware. Cost claims are simulation-derived, not measured output.
What the money actually buys
The round was split, and the split matters. A $375 million Series C at a $3.5 billion pre-money valuation was co-led by NEA, Andra Capital, Atreides Management, Valor Equity Partners, and Dylan Patel's SemiAnalysis Capital. A C-1 tranche of up to $500 million was anchored by NEA and Jim Clark, the co-founder of Silicon Graphics and Netscape. The money funds three specific things: the Asimov tapeout, a 2 MW-plus engineering data center and emulation platform, and the production ramp of Titan.
Positron is not selling simulations. Atlas, its first-generation system, is already deployed at Oracle Cloud Infrastructure, with more than 50 racks in production. Parasail, a partner, runs its own inference service on that capacity. Additional Atlas customers include Jump Trading and i3d.net.
Our focus now is to tape out Asimov, bring Titan to production, and scale manufacturing to meet the demand in front of us.— Mitesh Agrawal, CEO, Positron AI
That is the part of the pitch that has shipped. The part that has not is the headline efficiency number. Positron says simulations show a rack of its silicon processing up to 26 times as many tokens per dollar as Nvidia's GB300 NVL72, and roughly 5 times as many tokens per watt as Nvidia's Rubin. Both figures come from modeling, not from a production deployment. Treat them as a design target, not a benchmark.
Positron is betting that cheap, co-located LPDDR5X beats expensive HBM at serving long-context models.
The valuation is a 2027 wager: the product that justifies it is not in production yet.
The risk inside the pitch
The memory-first argument has a real hole in it. Skipping HBM caps peak bandwidth. For workloads that are bandwidth-bound rather than capacity-bound — high-throughput, short-context, batch serving — that tradeoff is a disadvantage, not an edge. Positron's answer is that long-context agentic work is the growth segment, and there its capacity advantage compounds.
The competitive field is also crowding fast. Nvidia is not standing still on inference efficiency. AMD is integrating hardwired inference silicon into its accelerator roadmap. And a generation of inference startups — Volantis, SiMa.ai, Rebellions, Euclyd — is attacking the same cost curve from different directions. Positron's differentiator is the memory choice, and it is a bet that only pays if the market's workloads shift the way it expects.
The tapeout is the next hard checkpoint. A design that works in emulation and fails in silicon would reset the story. Positron has to convert $875 million and a $5 billion valuation into a working chip on a 2027 timeline, against incumbents with shipping products and endless memory budgets.
Does memory-first inference win the serving layer by 2027?
Probability: 55% — the demand curve for long-context agentic workloads is real, but memory-first has to survive a first tapeout and prove its numbers outside simulation.
✅ Arguments for
Long-context agents reward capacity over peak bandwidth.
Confirmation criteria: Asimov tapes out on schedule and a named customer runs production inference on Titan in 2027.
❌ Arguments against
The 26x claim is simulation-derived, not shipped.
Falsification criteria: Asimov misses its tapeout window, or deployed Titan fails to beat HBM systems on tokens per dollar.
Development scenarios
🟢 Optimistic scenario (30%)
Consequences: Positron becomes the default alternative to HBM-centric serving and the $5B valuation looks early.
🟡 Base scenario (50%)
Consequences: A durable business, smaller than the round implies.
🔴 Pessimistic scenario (20%)
Consequences: Positron is re-priced toward its Atlas business, and the memory-first thesis recedes.