$800 million. That's the size of Together AI's Series C round announced July 1, 2026. The raise is one of the largest AI infrastructure raises of the year and the biggest bet yet that open-source models, not proprietary APIs, will deliver the next generation of enterprise AI.
The round values the San Francisco-based company at $8.3 billion, up from $3.3 billion just 16 months ago. Led by Aramco Ventures and joined by NVIDIA, Vista Equity Partners, General Catalyst, and Emergence Capital, the raise signals that the "neocloud" layer of AI infrastructure has graduated. Companies that rent GPU clusters purpose-built for open-weight models have become an institutional asset class.
Open-source AI infrastructure has reached production grade. Its annual bookings crossed $1.15 billion last quarter as enterprises abandoned closed-model APIs for cheaper open alternatives.
The neocloud sector is now a capital-intensive category of its own. Alongside its $800M, competitors Upscale AI raised $500M and TensorWave $350M in the last month alone, creating a new GPU-compute asset class between hyperscalers and bare-metal providers.
AI inference costs are collapsing 10x to 60x. Customers report saving 6x to 60x versus closed models as open-weight architectures matched or exceeded proprietary performance at a fraction of the cost.
The neocloud model: what Together AI actually does
The company builds what the industry calls a "neocloud": an infrastructure layer between hyperscalers (AWS, Azure, GCP) and bare-metal GPU rental. It provides on-demand access to NVIDIA H100 and Blackwell GPU clusters. The key differentiator is the software stack above the hardware: the platform is optimized end-to-end for open-weight models including DeepSeek V4, Meta's Llama 4, Mistral, Qwen, GLM, and Nemotron.
Customers don't just rent GPUs. They get a full-stack inference and training platform that, according to the company, delivers 6x to 60x cost savings compared to running the same workloads on closed-model APIs like OpenAI or Anthropic. Decagon, a customer, cut inference costs sixfold after migrating. Cursor and Cognition also run on its infrastructure.
Growth trajectory
Annual bookings surpassed $1.15 billion as open-source model usage tripled industry-wide. The company projects 50x infrastructure expansion over 5 years. · TechCrunch, July 2026
Why this round is different
The startup had already raised $102.5 million in Series A (2023) and $305 million in Series B (early 2025). The Series C at $8.3 billion marks a structural shift in who backs AI infrastructure. The round was led not by a traditional VC but by Aramco Ventures, the $7 billion venturing arm of Saudi Aramco. Sovereign wealth and strategic corporate capital, not just Sand Hill Road, are now the marginal dollars funding the compute layer.
The investor roster tells the same story. NVIDIA's participation is the chipmaker betting on the channel that distributes its hardware to the fastest-growing segment of demand. Vista Equity Partners, a $100 billion software-focused private equity firm, signals that neoclouds are no longer early-stage experiments but infrastructure assets with the recurring-revenue profile of mature enterprise software. Schneider Electric's SE Ventures invested because more efficient AI means less energy per workload, tying AI compute directly to the energy transition.
Enterprises are discovering that open-weight models have reached production parity with closed frontier models, and the economics favor the open stack by a wide enough margin that it now makes sense to build infrastructure specifically for it.
The neocloud sector becomes an asset class
The company is not alone. The neocloud sector has consolidated into a recognizable category over the last 90 days, with three of the largest rounds in AI infrastructure history closing within weeks of each other.
| Company | Round | Valuation | Lead investor |
|---|---|---|---|
| Together AI | ✔ $800M Series C | $8.3B | Aramco Ventures |
| Upscale AI | ◐ $500M Series A+ | $2.0B | Multiple |
| TensorWave | ◐ $350M Series B | $1.55B | AMD-aligned |
At a combined $1.65 billion in new capital across three companies in 30 days, the neocloud segment is absorbing more institutional capital than most growth-stage software verticals. The common thread: each is building GPU infrastructure optimized for a specific niche. It targets open-weight models; another focuses on enterprise SLAs for inference; a third differentiates on AMD hardware. The market is segmenting before it scales.
AI inference economics: the cost collapse powering the shift
The engine behind the company's growth, and the neocloud sector more broadly, is the extraordinary deflation in AI inference costs. Enterprise AI token costs fell 67% year-over-year in Q1 2026, according to AI.cc's infrastructure report analyzing 2.4 billion API calls across 8,000 accounts. Open-source and open-weight models now capture 38% of enterprise token volume, up from 11% a year earlier.
Three mechanisms drove the collapse. First, open-source model pricing created a new floor: DeepSeek V4-Flash launched at $0.14 per million input tokens, forcing broad repricing across the model ecosystem. Second, multi-model routing went from experimental to default: enterprises increased average model usage from 2.1 to 4.7 models per account in one year, routing simple tasks to cheap models and complex ones to frontier models. Third, aggregation-scale pricing from platforms like OpenRouter and AI.cc compounded discounts for high-volume customers.
Taken together, the result is that running a capable AI model in 2026 costs roughly 1,000x less per token than it did three years ago. That ranks as one of the fastest cost declines in computing history. The collapse created the demand that Together AI, Upscale AI, and TensorWave are racing to serve.
What happens to the market a year from now?
Its 50x capacity target implies a deliberate overshoot: building compute before demand materializes. The company that secures the most GPU allocation at favorable terms over the next 18 months will define the category. NVIDIA controls the supply; its investment is a natural hedge that also guarantees demand for its hardware.
Probability: 65 percent. The neocloud leader 18 months from now will be the one that signed the largest multi-year GPU commitment in 2026.
✅ Arguments for
+ Multi-model routing locks enterprises into platform switching costs
+ Inference demand is growing faster than training demand as deployed AI scales
Confirmation criteria: It announces a multi-year, multi-thousand-GPU deal with a hyperscaler or sovereign wealth fund.
❌ Arguments against
− Open-source model commoditization could compress margins faster than volume growth can offset
− The 500 MW compute commitment is capital-intensive and may not generate returns if demand growth slows
Disconfirmation criteria: A major hyperscaler launches a dedicated open-source AI cloud product with aggressive pricing.
Development scenarios
🟢 Optimistic scenario (30%)
Implications: Aramco Ventures' early bet on AI infrastructure becomes a template for sovereign wealth fund deployment in compute assets.
🟡 Base-case scenario (50%)
Implications: Neoclouds become the AWS of open-source AI: not the only option, but the default for workloads that need flexibility across models.
🔴 Pessimistic scenario (20%)
Implications: The infrastructure layer of AI converges back to the hyperscalers, and independent compute providers remain marginal.
GPU pricing trends for H100 and Blackwell: sustained price declines would validate the neocloud thesis; a price floor would signal hyperscaler pushback
Open-source model benchmark parity vs. GPT-5 and Claude 4: if open models maintain parity, the shift accelerates; if proprietary models pull ahead, the thesis weakens
Hyperscaler response: dedicated open-source AI products from AWS, Azure, or GCP with aggressive pricing would be the first sign of structural competition
Its next funding round: whether it is equity or debt, and whether trailing valuation increases or holds, will reveal how investors view the category 12 months from now