A 1.3-billion-parameter model trained on over a million hours of weather data now runs inside battery storage dispatch software in Poland. It predicts the atmosphere further ahead than the supercomputer standard, in minutes, on a fraction of the compute. It is a production configuration, and it is spreading.
Microsoft's Aurora, Google DeepMind's GenCast and NVIDIA's Earth-2 family all outperform the European Centre for Medium-Range Weather Forecasts (ECMWF) ensemble on renewable-relevant fields, at a fraction of the compute.
The open-weight shift is real, not announced. Aurora 1.5 shipped in July 2026 with 22 energy-relevant variables; grid operators and traders have already put Earth-2 into operation.
The economic tail is large. The International Renewable Energy Agency (IRENA) estimated that a 10% improvement in 24-hour wind forecasts could cut European grid balancing costs by €1.5–3 billion a year.
This is the forecasting layer of the grid changing hands. Traditional numerical weather prediction (NWP) runs physics equations on supercomputers and updates every few hours. The new class runs learned atmosphere models on GPU clusters and updates in minutes, with probabilistic output built in. The difference matters most where the grid already depends on weather: wind, solar and the battery storage that smooths both.
None of this needs to be believed on paper. The models are open, the deployments are named, and the numbers are public.
What the models now beat
The benchmark that matters is not generic forecast skill. It is skill on the fields that drive generation forecasts: hub-height wind, irradiance, ramp events and extremes. DeepMind's GenCast, published in Nature in December 2024, beat the ECMWF ensemble system on 97.2% of its test cases while producing a 15-day forecast in 8 minutes on a single Cloud TPU. The ECMWF system needs hours on a supercomputer with tens of thousands of processors.
Microsoft's Aurora is not a time-series transformer tuned for one market. It is a 3D Swin Transformer trained on over a million hours of weather and climate data, predicting the atmospheric state six hours ahead at 0.25-degree resolution, then rolled forward. It was fine-tuned for air quality, ocean waves, tropical cyclone tracks and high-resolution weather, each at orders of magnitude smaller compute than the dedicated systems.
NVIDIA's Earth-2 family, launched in January 2026, pushed the same logic toward open tooling. Earth-2 Medium Range, built on an architecture called Atlas, covers 15 days across more than 70 weather variables. Earth-2 Nowcasting uses generative models to turn country-scale forecasts into kilometer-resolution, zero-to-six-hour storm predictions in minutes. Earth-2 Global Data Assimilation produces the initial conditions a forecast needs in seconds on GPUs, where the old path took hours on supercomputers.
The headline numbers, for the record:
Aurora's trained size
Trained on over a million hours of weather and climate data; predicts the atmosphere at 0.25° resolution. arXiv, 2024
AI over the ensemble standard
15-day forecast in 8 minutes on one TPU; the ECMWF ensemble takes hours on a supercomputer. Nature, 2024
Earth-2 Medium Range reach
Open model family used by grid operators and traders; CorrDiff downscales up to 500x faster than traditional methods. NVIDIA, 2026
The gap that pays
Better wind and solar forecasts convert into money through balancing costs. IRENA's estimate puts the value of a 10% forecast improvement for 24-hour wind at €1.5–3 billion a year on the European grid alone. That figure predates the current models, which is why the operator-side interest is not curiosity.
Storage is where the effect shows up first. A battery decides hours in advance whether to charge, hold or discharge, and that decision is only as good as the price and generation forecasts feeding it. Foundation models change the economics of that layer twice. They output probabilistic quantiles natively, which is exactly what risk-aware dispatch wants, and they do zero-shot, so a new site does not need months of bespoke modeling.
A concrete case is documented from Poland, where a BESS operator runs a composite stack: a time-series foundation model with a residual model on top for fundamental drivers like gas and carbon prices. Across six sites in 2026, day-ahead forecasts averaged 12.3% mean absolute percentage error, versus 14.8% for the pure gradient-boosting baseline of 2024. The difference is structural: two points of improvement in the metric that sets dispatch revenue.
The same logic applies to solar. Solcast, the incumbent standard for solar irradiance, fused satellite and cloud-motion data with a weather model and posted the lowest error in an EPRI benchmark of large US plants. Foundation models add the long-horizon component that irradiance vendors are weak at.
This is the direction, and it is consistent across the market.
The cost curve that decides adoption
The structural advantage is compute, and it is measurable. A 50-site utility portfolio runs foundation-model inference for roughly $800–1,500 per month in cloud GPU compute, plus another $200–500 for storage and orchestration. For that, the operator gets probabilistic forecasts across every site without building a single per-site model.
The comparison to the old way is the point. A per-site LSTM or gradient-boosting model needs clean historical data per asset, a modeling cycle per asset, and rework whenever the fleet changes. A foundation model arrives pretrained on hundreds of billions of time points, which no single site could ever match, and generalizes zero-shot. The composite stack used in the Polish case beats both pure foundation-model and pure fundamental forecasts by 2–4 mean absolute percentage points in day-ahead error.
This is why the vendor ecosystem is forming so fast. Solcast covers the short-horizon solar layer, specialist firms sell day-ahead and intraday prices, and the foundation models sit underneath as the common input layer. The integration work, not the model training, is where the value migrates.
Open weights met grid operations
The 2026 shift is not another paper. It is deployment. NVIDIA's Earth-2 launch listed named energy users: TotalEnergies is evaluating Nowcasting for short-term risk; Eni is testing downscaling for probabilistic weather and gas-demand forecasts; GCL, a Chinese solar producer, is running Earth-2 in operation for its photovoltaic prediction; Southwest Power Pool, with Hitachi, is using Nowcasting and FourCastNet3 for intraday and day-ahead wind forecasting across its footprint.
Jua, a Zurich startup founded in 2022, built its whole product around this. Its physics-constrained model targets energy traders directly, and it claims to outperform ECMWF on the fields trading desks trade on. Jua has raised roughly $27 million from investors including Ananda Impact Ventures and Future Energy Ventures.
Microsoft extended the open path in July 2026 with Aurora 1.5. It added 22 weather variables relevant to energy, agriculture and transport, plus hourly temporal resolution and probabilistic ensemble forecasting, and released it open-source with checkpoints on Hugging Face. The rationale from the energy side is blunt. BKW, the Swiss utility in the announcement, said AI forecasts support its ambition to run a renewable-based system where generation is inherently weather-dependent, and to manage that variability with more confidence.
What makes this durable is the open-weights property. A utility can fine-tune, run on its own infrastructure, and keep sovereign control of a national-security-adjacent capability. NVIDIA's Mike Pritchard made the point explicitly at the January launch: weather is a national security issue, and sovereignty and weather are inseparable.
Three open models, one job
| Dimension | Aurora 1.5 | GenCast | Earth-2 Medium Range |
|---|---|---|---|
| Vendor | Microsoft | Google DeepMind | NVIDIA |
| Core skill | ✔ Earth-system foundation, energy variables | ✔ Ensemble probability, wind | ✔ 15-day, 70+ variables, open tooling |
| Update cadence | Hourly (1.5) | 6-hour steps, 15-day horizon | Medium-range + nowcasting + assimilation |
| Open weights | ✔ GitHub + Hugging Face | ✔ code + weights | ✔ fully open family |
| Named energy users | BKW, Polish BESS ops | Wind-power trials | TotalEnergies, Eni, GCL, Southwest Power Pool |
None of the three is a full replacement for a national meteorological service. They are the layer that turns atmosphere forecasts into operationally usable generation and price forecasts, and they are converging on the same open-weights model that dominates large language models: a general pretrained core, specialized by fine-tuning and integration.
Signals to track
Whether national forecast agencies (ECMWF first) license or build AI-native successors rather than guard the NWP stack.
Whether ISO-provided forecasts start carrying AI-model outputs, which would reset the baseline for every market participant.
Whether storage software vendors make foundation-model forecasting the default EMS layer, as the Polish case suggests.
Whether the EU AI Act's high-risk classification for grid-management AI creates a market for regulator-certified forecasting models.
The grid's forecasting layer is being rebuilt on learned atmosphere models, and the building blocks are open. The models beat the old standard, the operators are named, and the money is measurable. The open question is timing: how fast the incumbents in the middle, the ones selling physics-equation forecasts into trading desks, get rebuilt or get displaced.