The cheapest GPU by the hour isn't the cheapest for your workload. And the most expensive one might save you money.
But egress fees, idle time, and architectural lock-in can erase those savings entirely for workloads that don't match the pricing model.
The provider that minimizes total cost depends on whether you train, serve inference, or both, and whether your workload tolerates interruption.
In early 2026, the big five hyperscalers committed roughly $660 billion in capex, most of it directed at AI infrastructure. AWS alone plans to deploy over a million NVIDIA GPUs this year. Meta raised its 2026 capex guide to $125–145 billion, and Alphabet is spending $180–190 billion. The scale is unprecedented.
But for the teams actually running AI workloads (startups, mid-market engineering orgs, enterprise AI groups), the hyperscaler pricing menu looks different from what the headline numbers suggest. A single H100 on p5 instances costs $6.88/hr if you can get one. On GCP it runs $10.98/hr. Azure charges $12.29/hr. Meanwhile, neoclouds (specialized GPU providers like CoreWeave, Lambda Labs, and Nebius that build no-frills infrastructure for compute workloads) offer the same H100 for $2–4/hr, no contract required and often with free data egress.
Compute Layer: Neocloud vs Hyperscaler
The gap isn't an edge case. Saturn Cloud's June 2026 report on 17 GPU cloud providers found that self-service H100 pricing spans roughly $1.80–6.16/hr depending on provider and commitment level, compared to $6.88/hr, $10.98/hr on GCP, and $12.29/hr on Azure. That is a three-to-six times difference for the identical NVIDIA GPU hardware. CoreWeave charges no egress for any data transfer. Crusoe charges none either. The hyperscalers charge $0.08–0.12/GB for data leaving their network, a cost that can add 20–40% to a monthly bill for data-heavy pipelines.
But the neocloud story is not uniformly cheap. Several raised published on-demand H100 rates in early 2026: Lambda went from $2.99 to $3.99–4.29/hr; Verda (formerly DataCrunch) moved from $2.29 to $3.25/hr. Nebius held steady at $2.95/hr. The blanket assumption that neoclouds undercut hyperscalers is now provider-specific. Check before committing, and verify that the published rate reflects actual availability rather than a theoretical list price with no real inventory behind it.
| Provider | H100 On-Demand | H100 Spot | Egress | Billing |
|---|---|---|---|---|
| CoreWeave | ~$3.50/hr | N/A | Free | Per-hour |
| Lambda Labs | $3.99–4.29/hr | N/A | Included | Per-hour |
| Nebius | $2.95/hr | N/A | Free | Per-minute |
| Spheron | $2.50/hr | $1.03/hr | Included | Per-minute |
| AWS (p5) | $6.88/hr | ~$3.83/hr | $0.09/GB | Per-second |
| GCP | $10.98/hr | N/A | $0.12/GB | Per-second |
| Azure | $12.29/hr | N/A | $0.087/GB | Per-hour |
Spot instances widen the gap further. Spheron lists H100 SXM5 spot at $1.03/hr, 59% below its own on-demand rate and roughly 85% below hyperscaler spot pricing. Spot instances can be reclaimed with 30 seconds to 2 minutes of notice. For batch training jobs with checkpointing, that's tolerable. For production inference, it isn't.
Inference vs Training: Two Different Cost Curves
The pricing calculus shifts depending on workload type. Training is throughput-bound. You care about total FLOP-hours and interconnect bandwidth. Inference is latency-bound and memory-bound. You care about tokens per second and VRAM capacity per dollar. A provider that offers the lowest H100 price may be a poor fit for inference workloads if its networking adds latency or its instance types don't match the memory profile of your model.
For training, the neocloud advantage is clear: the same H100 silicon costs 60–80% less, and NVLink-based multi-GPU configurations are standard. CoreWeave and Lambda both offer InfiniBand clusters for distributed training at a fraction of hyperscaler rates. Meta's own cloud launch (which we covered here in July) signals that even hyperscaler-native workloads are starting to evaluate external GPU capacity as a cost-management lever.
A 70B-parameter model served at production scale needs very different infrastructure than a fine-tuning job for a small 7B model. For inference, the math is different. Hyperscalers offer integrated serving stacks (Amazon SageMaker, Google Vertex AI, Azure ML) that reduce operational overhead. A team serving a 70B-parameter model at scale might find that the hyperscaler's managed inference platform, despite higher GPU cost, delivers better total cost when engineering time and operational complexity are factored in. A startup running a single production model on SageMaker at $10/hr for H100 inference avoids the DevOps overhead of self-hosting vLLM on a neocloud at $3/hr. The $7/hr premium buys managed scaling, monitoring, and automatic failover that would cost more in engineering hours to replicate.
The AMD MI300X, available at $2.50–3.00/hr with 192 GB of HBM3 memory, adds another important variable. It is 20% cheaper than H100-class NVIDIA silicon, with 5.2 TB/s memory bandwidth that gives it a 15–20% inference speed advantage on memory-bound LLMs. For teams running large-scale inference on models that are memory-bandwidth-limited rather than compute-limited (which describes most modern LLM deployments), the MI300X presents a rare case where a hyperscaler's GPU offering is price-competitive with neocloud NVIDIA instances.
Hidden Costs Beyond the Hourly Rate
The hourly GPU rate is the visible number. The hidden costs determine the actual bill.
Egress is the largest. Hyperscalers charge $0.08–0.12/GB for outbound data. For a team running daily inference pipelines serving thousands of requests, egress can rival the compute cost. Most neoclouds charge zero egress. A team moving 50 TB/month in data would pay $4,000–6,000/month in egress fees at hyperscaler rates, effectively a 20–40% surcharge on GPU compute.
Idle GPU time is second. Teams often spin up GPU instances and forget to shut them down after a job completes. A single idle H100 on AWS costs $4,950/month. On Spheron or Nebius, the same idle instance runs $1,800–2,100/month. The savings compound across fleets. Per-minute billing (Nebius, Spheron) further reduces waste for short-duration workloads like hyperparameter sweeps and data preprocessing.
Egress: Hyperscalers add $0.08–0.12/GB for outbound data. At 50 TB/month, that's an extra $4,000–6,000.
Idle instances: A single idle H100 costs $4,950/month on AWS vs $1,800–2,100 on neoclouds.
Commitment lock-in: Hyperscaler reserved instances (1–3 year) discount GPU cost by 20–40% but freeze hardware selection. Next-generation GPUs arrive within the contract term, making committed capacity a depreciation risk.
The Market Forces Driving the Gap
The neocloud pricing advantage is not a temporary startup subsidy or a promotional discount. It reflects a fundamentally different cost structure. Hyperscalers build general-purpose cloud regions with hundreds of services (object storage, databases, serverless compute, CDN, identity management), each with its own engineering team, compliance burden, and margin target. The GPU instances sit inside this sprawling machine, carrying overhead from services that AI teams may never use.
Neoclouds build one thing: GPU compute. CoreWeave operates a single-product infrastructure with no database-as-a-service, no serverless functions, no content delivery network. Every dollar of engineering spend goes to GPU networking, storage, and orchestration. The result is a cost base that naturally runs 40–70% below hyperscaler pricing on equivalent hardware.
The market has noticed. By July 2026, the North America GPU-as-a-Service market was projected to reach $15.34 billion by 2030, growing at 26–28% CAGR, according to Research and Markets. Specialized GPU cloud providers now compete directly with hyperscalers on procurement cycles that were previously captive to AWS, GCP, or Azure, and they are winning on price.
Provider Deep-Dive: Three Neocloud Strategies
CoreWeave, the largest pure-play neocloud, went public in early 2026 with a $99.4 billion revenue backlog. Its strategy is hyperscaler-grade infrastructure without the hyperscaler margin structure. It offers InfiniBand clusters, VAST storage integration, and NVIDIA's full GPU stack from H100 to B200. It charges roughly $3.50/hr for H100 on-demand, with no egress fees. The trade-off: its pricing model rewards reserved commitments, and on-demand availability can be tight during peak demand.
Lambda Labs built its reputation serving academic and research teams. Its on-demand H100 PCIe runs $3.99–4.29/hr, up from $2.99/hr in late 2025, a 33–43% increase that reflects tightening supply. Lambda operates its own fleet, which means when inventory fills up, customers wait. The company's 3-year reserved contracts bring H100 cost to $2.43/hr, but locking in for three GPU generations in a market where NVIDIA refreshes architecture annually carries depreciation risk.
Nebius, the European neocloud, holds at $2.95/hr for H100 on-demand with per-minute billing, the most aggressive pricing among the three. Its $17.4 billion Microsoft contract, signed in late 2025, signals that even hyperscalers are outsourcing GPU capacity to neoclouds. Nebius bills per minute rather than per hour, which for short-duration workloads (preprocessing, hyperparameter sweeps) can cut costs by 30–50% vs per-hour billing.
The AMD Factor
AMD's MI300X changes the neocloud pricing calculus in ways that the H100-centered comparison misses. With 192 GB of HBM3 memory and 5.2 TB/s bandwidth, the MI300X matches or exceeds the H100 on memory-bound inference workloads and costs 20% less on Azure. A team running LLM inference at scale gets 15–20% higher throughput on the MI300X thanks to its memory bandwidth advantage, at a lower per-hour rate.
NVIDIA's CUDA ecosystem, TensorRT, and NCCL libraries remain the default for production AI workloads. AMD's ROCm has improved rapidly, supporting PyTorch, vLLM, and the major inference frameworks, but production deployments still require more engineering time to match NVIDIA's out-of-box performance. For cost-conscious teams willing to invest in software optimization, the MI300X is the best alternative to the H100 on the market today.
AMD's next-generation MI400 series, expected in late 2026, targets both training and inference with a unified architecture. If it delivers on performance parity, the GPU cloud pricing market could shift from a neocloud-vs-hyperscaler binary to a three-way competition where AMD introduces real price discipline across both tiers.
When Hyperscalers Still Win
Hyperscalers remain the right choice in three scenarios. First, deep integration with cloud-native services: if your stack runs on S3, BigQuery, and IAM policies, the switching cost of moving data out exceeds the GPU savings. Second, compliance requirements: financial services and healthcare orgs often need certifications that neoclouds haven't prioritised. Third, production inference at scale where spot interruption is unacceptable and managed ML platforms reduce engineering overhead.
The neocloud trade-off is operational: lower GPU cost, free egress, flexible billing, but less platform integration, fewer certifications, and a self-service model that expects the user to bring their own orchestration. For teams that already run Kubernetes and manage their own inference stacks, the neocloud is the better financial decision. For teams that want to offload infrastructure management, the hyperscaler premium buys convenience.
Decision Framework
For batch training with checkpointing: neocloud spot instances deliver 60–85% savings over hyperscaler on-demand with manageable interruption risk.
For production inference: evaluate total cost including egress, orchestration, and engineering time. Neocloud on-demand is 40–70% cheaper on GPU cost alone; managed hyperscaler inference may still be cheaper overall for small to mid-scale deployments where the overhead of self-managed serving outweighs the GPU savings.
For multi-node training requiring InfiniBand: neoclouds offer competitive clusters. CoreWeave and Lambda both support RDMA fabrics at rates well below hyperscaler equivalent configurations. An 8x H100 training node that costs roughly $55/hr on AWS (p5.48xlarge) can be provisioned for approximately $28/hr, a 49% saving that compounds over weeks-long training runs.
For compliance-constrained workloads: hyperscalers remain the default. The neocloud certification market is developing but not yet comparable for regulated industries. Financial services firms bound by SOC 2 Type II and PCI DSS, or healthcare organisations subject to HIPAA, will find few neoclouds that document these certifications as thoroughly as AWS or Azure do.
Spot vs on-demand is not a fixed decision. A team that runs nightly batch inference jobs can use spot at $1.03/hr with checkpointing and tolerate the occasional interruption. The same team running a 24/7 production API must use on-demand or reserved capacity. The gap between the two ($1.03/hr vs $2.50/hr on Spheron, $3.83/hr vs $6.88/hr) defines the premium that reliability demands. A mid-size AI team running a mixed workload (70% batch, 30% production) can cut total GPU spend by roughly 40% by routing batch jobs to spot and reserving on-demand for serving only, rather than provisioning all capacity at on-demand rates.
The 2026 GPU cloud market is not a simple neocloud-vs-hyperscaler binary. The gap is real, 40–85% on raw GPU cost, but the right provider depends on workload shape, data gravity, and operational tolerance. The teams that get this right will be the ones that calculate total cost per token or per training step, not cost per GPU-hour. A team running 10,000 hours of H100 training per quarter at $2.50/hr on a neocloud instead of $6.88/hr on AWS saves approximately $131,000 per quarter on compute alone. The same team choosing a hyperscaler for its managed inference platform might pay more per GPU but save twice that in engineering time. The answer depends on knowing which cost layer your organisation actually spends on.
Two developments could reshape this market in the next 12 months. If hyperscalers cut egress pricing (a plausible response to neocloud competition), the hidden-cost advantage that neoclouds currently enjoy would shrink. If AMD's MI400 series delivers real training parity, the GPU pricing floor drops across both tiers. Either scenario rewards teams that stay flexible: those that avoid long-term commitments, benchmark workloads across providers, and treat infrastructure procurement as an active optimisation problem rather than a one-time vendor choice.
Hyperscaler pricing responses: if hyperscalers narrow the GPU pricing gap, the neocloud thesis weakens.
AMD MI400-series availability: a credible second source of training silicon would reshape the pricing market across both tiers.
Egress pricing changes: neoclouds currently use free egress as a differentiator; hyperscaler egress cuts would remove a major cost advantage.