The AI build-out keeps running into the same wall, and the wall is not made of silicon. A gigawatt-scale data center costs $9–19 billion and takes years to permit, wire and energize. The models it is meant to serve now change on a quarterly cadence. Runware, an inference company founded in 2023 with offices in London and San Francisco, has spent two years making the opposite bet: one megawatt of purpose-built compute inside a 20-foot shipping container, placed wherever power already exists.
In August, it shipped.
The power wall
Data centers were designed for websites and databases. Inference behaves differently from a website: it does not care where it runs, only what it costs per token. Yet the industry keeps buying the same overbuilt shell — backup systems, redundant power, cooling for hardware that sits idle much of the day.
The arithmetic has stopped working. Roughly a third of a typical data center's electricity goes to cooling, power conversion and building systems rather than to compute. Global data-center demand is projected to roughly double by 2030. Grid interconnection queues in the US and Europe stretch for years, so an operator can raise the capital and still wait half a decade to switch on.
As we wrote in September, Crusoe's $3B round put a $30B price on the same thesis — energy-first AI infrastructure. Crusoe has raised again since and has started building small factory-made facilities for inference. Runware's container takes the thesis to its physical limit.
Facilities cost runs up to 100× lower than a gigawatt-scale build, at 30–90% lower inference cost per GPU-hour.
160 sites are booked, the first 10,000 nodes arrive through 2026, and the 2027 target of more than 1 GW roughly equals all US capacity under construction for 2026.
One megawatt, twelve hundred GPUs, no water
A Sonic Inference Pod carries one megawatt of IT load inside a standard 20-foot container, with a chiller about the size of the container mounted on top. Inside sit around 1,200 GPUs in nodes of two to eight, built on custom servers with no cases, custom racks, in-house PCIe switching, high-frequency CPUs and local NVMe storage. Nvidia's RTX PRO 6000 does most of the work; B200 and B300 chips handle the larger jobs.
Runware Sonic Inference Pod density
About 1,200 GPUs and 1 MW of IT load in one 20-foot container. Runware, 2026
Cooling is closed-loop and liquid, and Runware says the pod consumes no water — against evaporative systems that, across the sector, draw down an estimated 560 billion liters a year. The unit arrives on a truck and needs three things: ground, power and a network connection. Build time runs about three weeks, installation about a day.
Software does the rest. Every pod joins one distributed inference network that reroutes a request when a unit drops offline, trading some locality for resilience without overbuilding any single site.
The economics: cheaper to build, and to run
The claim is that inference should be priced by throughput rather than by real estate. A pod carries none of the oversized shell, so facilities cost falls by as much as 100× against a gigawatt-scale build. On the operating side, Runware puts inference at 30–90% below typical providers per GPU-hour — 50% or more for most workloads — at roughly twice the throughput.
| Parameter | Traditional data center | Sonic Inference Pod |
|---|---|---|
| Build time | ✗ Years to permit and connect | ✔ ~3 weeks build, ~1 day install |
| Facilities cost per GW | ✗ $9–19B | ✔ Up to 100× lower |
| Cooling | ◐ Evaporative, ≈560bn L/yr | ✔ Closed-loop, water-free |
| Siting | ✗ Stuck in grid queue | ✔ Wherever power exists |
| Inference cost per GPU-hour | Baseline | ✔ 30–90% lower |
Runware, 2026; build-cost range per company figures.
What still has to be proven
The plan is ambitious on paper. Europe is live, US West is underway, and Runware says it has 160 locations booked, with the first 10,000 nodes arriving through 2026 and more than one gigawatt targeted for 2027. That 2027 figure is roughly equal to all US data-center capacity under active construction for 2026 delivery.
Cost-per-GPU-hour figures are self-reported and workload-dependent. Siting at solar parks chases cheap power, not a cheap grid — intermittent generation still needs firming. Each of the 160 sites carries its own permit, land and connection problem.
None of that turns the container into a gimmick. It makes the box a test.
If inference really is indifferent to location, the winners of the next compute cycle will not be the operators with the biggest campuses. They will be the ones who can follow the electrons. Runware is wagering that a 20-foot container follows them fastest.