What if the hardest constraint in AI infrastructure isn't how fast a chip computes, but how fast it can feed that compute? Positron AI just raised $875 million on the bet that it is, and that the memory most vendors treat as a commodity is the part worth redesigning around.
Its next-generation Asimov chip drops HBM for commodity LPDDR5X memory and claims more than 90% realized memory bandwidth, against under 30% for typical GPUs.
The cash funds a tapeout at the end of 2026 and a production ramp in the second half of 2027, so the thesis stays unproven in silicon for now.
The financing closed in two tranches: a $375 million Series C at a $3.5 billion pre-money valuation, and a Series C-1 of up to $500 million anchored by NEA and Jim Clark, the Netscape co-founder. Evertiq reports NEA, Atreides Management, Valor Equity Partners, Andra Capital and Dylan Patel's SemiAnalysis Capital as co-leads. Qatar's sovereign wealth fund returned after backing the company in February.
Positron is a Reno, Nevada company founded in 2023, with 86 employees and a first-generation inference appliance, Atlas, already shipping. More than 50 racks run at Oracle Cloud Infrastructure, serving customers that include the trading firm Jump Trading and the hosting provider i3d.net.
TIMELINE: Positron AI
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Feb 2025 โโโโ Feb 2026 โโโโ Sep 2026 โโโโ Late 2026 โโโโ H2 2027
๐ฑ ๐ฆ ๐ฐ ๐ญ โ
Seed $230M B $875M C Asimov Titan
$23.5M at $1B+ at $5B tapeout N3P production
Funding rounds and product milestones, Positron AI
Inference turned into a memory problem
Training a model is a capital event. Serving it is a recurring one. Every assistant and agent issues inference calls, and each call repeats the same mechanical step: pull the model's weights and the conversation's context out of memory, run the arithmetic, write the result back. Models now hold billions of parameters and context windows that stretch into the millions of tokens. Both have to move through memory on nearly every token.
That makes bandwidth the binding constraint. Positron's pitch, as reported by SiliconANGLE, is that most accelerators leave their high-bandwidth memory badly underused. The company puts a typical GPU at under 30% of available HBM bandwidth, against more than 90% of LPDDR5X throughput on its own hardware. The claim is about utilization, not peak speed.
Power sharpens the same point. Tokens per watt, not tokens per second, sets the operating cost of an inference fleet, and bandwidth that goes unused is energy spent for nothing. A chip that reaches more of its memory's throughput can serve more tokens on the same electricity.
Asimov's claimed memory utilization
Company figure, against under 30% for typical HBM-equipped GPUs. ยท Positron AI / SiliconANGLE, 2026
The memory choice that carries the risk
HBM is the expensive part of an AI accelerator. It is scarce, it needs advanced packaging capacity, and its supply sits with a handful of manufacturers. LPDDR5X, the memory used in phones and laptops, is made at volume by several suppliers at a fraction of the cost.
Asimov is built around that trade. According to the company's Asimov specification page, each chip carries 864 GB to 2.3 TB of LPDDR5X, moves data at 2.76 TB per second, draws about 400 watts, and runs air-cooled. The company claims roughly six times the memory capacity per chip of an HBM design at a lower system cost.
Dropping HBM also means dropping CoWoS, the advanced packaging process that binds memory to logic and that has been one of the tightest links in the AI supply chain. That is a supply decision as much as a technical one, and it is why the design can commit to LPDDR5X volumes the round explicitly funds.
| Parameter | LPDDR5X (Asimov) | HBM (typical GPU) |
|---|---|---|
| Memory per chip | โ 864 GB โ 2.3 TB | โ Smaller |
| Realized bandwidth | โ Above 90% | โ Under 30% |
| Advanced packaging | โ Not required | โ CoWoS-class flow |
| Peak bandwidth | โ 2.76 TB/s | โ Higher on paper |
| Supply base | โ Commodity, multi-source | โ Constrained |
Asimov memory bandwidth
Realizable throughput per chip on commodity LPDDR5X. ยท Positron AI, 2026
LPDDR5X trails HBM on paper bandwidth. Its case rests on closing that gap through utilization, a claim about real workloads rather than a spec-sheet win.
What a $5 billion price is underwriting
Seven months after reaching a $1 billion valuation, the company is priced at $5 billion. The round drew a mix of crossover funds and strategic names: NEA, Atreides, Valor, Andra Capital, Cisco Investments, Naver Ventures and Qatar's sovereign wealth fund.
Two co-leads sit close to the company's technical story. Dylan Patel's SemiAnalysis Capital is tied to the benchmarking community that scrutinizes inference claims. Jim Clark, who founded Silicon Graphics and Netscape, anchored the second tranche. As part of the deal, NEA's Forest Baskett, Atreides' Gavin Baker, Jim Clark's Thomas Jermoluk and Patel join the board.
What the price underwrites is a thesis, not a balance sheet. Investors are paying for the idea that the cost of serving a token can be reset by changing which memory a chip uses, and by doing so outside the constricted HBM chain.
Positron's September funding round
Split into a $375M Series C and a Series C-1 of up to $500M, at a $5B post-money valuation. ยท Reuters, 2026
Unproven silicon, and a 2027 clock
Two things stand between the thesis and the market. The first is physics that has not shipped. Positron states that Asimov's performance figures come from cycle-accurate simulations, not working silicon. Tapeout on TSMC's N3P process is due at the end of 2026, with production in the second half of 2027. A simulation is a planning tool, not a product.
The second is time. By the 2027 ramp, the incumbent accelerator roadmap will have moved again, and any easing of HBM supply would soften the scarcity its architecture is built to avoid. Execution risk compounds both. A defect in the first tapeout, or a slip in the LPDDR5X supply commitments the round is meant to lock in, pushes the whole story further out.
A memory-first architecture is only as good as the silicon that proves it. Until Asimov tapes out and runs a paying workload, the 90% utilization claim is a simulation, and the market is valuing a roadmap.
Why it matters beyond one chipmaker
As we wrote in September, the inference economy has a fixed shape: more requests, longer contexts, tighter power budgets. If Positron is right, memory selection stops being a procurement detail and becomes an architectural decision. A commodity memory that is cheap and well-utilized can beat a scarce premium part that sits half-idle.
The scoreboard is simple to track. Asimov's tapeout, the first paying Titan deployments, and whether the claimed utilization holds on real customer workloads will settle the argument faster than any benchmark slide. When a startup can raise $875 million on a memory architecture claim alone, the bottleneck the market priced as compute has quietly moved.