> ## Content Index
> Fetch the complete content index at: https://nexi.fund/llms.txt
> Use this file to discover other available public pages before exploring further.

# Arc Institute Virtual Cell AI
- URL: https://nexi.fund/arc-institute-virtual-cell-ai-2026/
- Published: 2026-10-07T10:00:07.000Z
- Updated: 2026-10-07T10:00:07.000Z
- Description: Virtual cell AI from Arc Institute could reshape biological research and drug discovery with predictive models.
- Author: Nexi.fund Labs
- Tags: Biotech & Health, #mode-1, #track-A, #hook-thesis

Arc Institute trained a model on single-cell observations from nearly 170 million cells — and on more than 100 million perturbation readouts. No rival in this corner of biology has trained on more.

State, released June 23, 2025, predicts how stem, cancer, and immune cells rewire their transcriptomes after drugs, cytokines, and gene knockouts. Arc reports it beats every published method on that specific task. That part is settled.

🎯

**Three things to take away**  
  
Data scale is the product. State trains on 167M observational cells plus 100M+ perturbation cells across 70 lines; the Arc Virtual Cell Atlas now clears 300M cells and targets over a billion experiments.  
  
Benchmarks build the moat. Two Arc-run Virtual Cell Challenges, with $100,000 prizes and NVIDIA on the sponsor list, set the yardstick every competing model is scored against.  
  
Nonprofit structure is an edge. No equity claim sits on State, yet target validation happens in the open — before the private market pays for it. 

## What State actually is

State is two transformer models. The Embedding component, trained on 167 million observational cells, builds a representation of a single cell. The Transition component, trained on more than 100 million perturbation cells across 70 cell lines, predicts how that representation shifts when a gene is knocked down or a small molecule lands on it.

The unit of prediction is a set of cells, not a single row.

Arc's evaluation claims are narrow but testable: State generalises to cell contexts it has never seen perturbed, which is the property a lab cares about when it plans the next experiment. The description ran in Cell (Adduri et al., 2026). The weights are public under a noncommercial licence.

100M+ perturbed cells 

#### Perturbation training data

Single-cell perturbation readouts (Tahoe-100M, Parse-PMBC, Replogle-Nadig) behind State's transition model. · *Arc Institute, 2025*

## The dataset pipeline is the moat

Models get copied. Data pipes do not. Arc's Virtual Cell Atlas, launched February 25, 2025, opened with more than 300 million single cells, uniformly reprocessed to remove analytical artifacts. scBaseCount, Arc's agentic curation framework, collates every public single-cell dataset and re-processes it on one standard.

Tahoe-100M, contributed by Tahoe Therapeutics, is part of the same release. On the sequencing side, Arc brought in 10x Genomics and Ultima Genomics in April 2025 to push per-cell cost down and throughput up. The 2026 Challenge data, generated on Ultima's UG100 platform with 10x Flex profiling, is the first visible output of that pipeline.

The investment lesson: dataset access decides who competes. Arc publishes its data openly, so rivals can train on the same rows. They cannot generate more rows at the same cadence — that capacity sits inside Arc's Technology Centers, against a stated target of over a billion perturbation experiments.

## The benchmark is the business layer

On June 26, 2025, Arc launched the Virtual Cell Challenge inside a Cell commentary: a $100,000 prize for the model that best predicts how cells respond to genetic perturbations, sponsored by NVIDIA, 10x Genomics, and Ultima Genomics.

Year two raised the bar. The 2026 Challenge, open since August 20, 2026, drops the training crutch: teams get six cell lines they have never seen perturbed and must predict CRISPRi knockdown responses from unperturbed control profiles. Three lines run the live leaderboard; three are held back for the final test. Same prize, stiffer test.

| Parameter        | 2025 Challenge                           | 2026 Challenge                            |
| ---------------- | ---------------------------------------- | ----------------------------------------- |
| **Task**         | Predict single-gene perturbation effects | Zero-shot prediction in unseen cell lines |
| **Data**         | Perturbation benchmark datasets          | CRISPRi knockdowns in six cell lines      |
| **Training set** | Provided                                 | None — control profiles only              |
| **Prize**        | $100,000                                 | $100,000                                  |

Arc Institute — Virtual Cell Challenge 2025 / 2026

Every lab that submits adopts Arc's benchmark as the yardstick for its own work. NVIDIA framed the sponsorship bluntly: the competition exists to empower builders of foundational models that predict how genetic perturbations change cells. The structure is familiar from computer vision — one referee, one number, everyone compared against it.

## Where the ceiling is

Three limits decide how far this goes. Generalisation: the zero-shot round is the honest test, and no model has passed it yet. Validation: an in-silico hit still has to survive the wet lab, then the clinic — two places where model performance routinely evaporates. Commercial wind: State's licence is noncommercial, so profitable use means rebuilding the data pipeline privately.

Nonprofit distribution works for science. Whether it works for revenue is a different question.

Biology has a long record of models that memorised batch effects instead of learning mechanism. That is precisely why the sector is sceptical and why an honest public benchmark is the cheapest insurance against the failure mode.

## Will a virtual cell redirect the money in drug discovery?

🔮

**By the end of 2028, at least one early-stage biotech financing will name a virtual-cell screening workflow validated against Arc's benchmark line as the basis for its lead-discovery engine.**  
  
Probability: 60% — the challenge series gives investors a repeatable, third-party frame for underwriting target validation before wet-lab scale-up. 

#### ✅ Arguments for a repricing

Open benchmarks shrink the diligence cost of validating a target before a financing round.  
  
A billion-cell data pipeline compounds faster than any single company's private dataset.  
  
No licence overhang on research use keeps academic spinouts inside the ecosystem.  
  
Endowment plus Audacious Project funding removes the early commercial pressure that distorts other foundation-model projects.  
  
**Confirmation criteria:** a Series A/B filing names virtual-cell screening; a third model beats State on the 2026 zero-shot task. 

#### ❌ Arguments for a lukewarm market

No revenue line exists for the model itself.  
  
The noncommercial licence pushes commercial teams to rebuild data pipelines privately.  
  
Across-patient generalisation is unproven — single-cell markers may not transfer to tissue.  
  
The foundation-model hype cycle has already burned wet-lab budgets once.  
  
**Disconfirmation criteria:** the zero-shot round shows no gain over simple baselines; a private, wet-lab-validated rival reaches the clinic pipeline first. 

## Three scenarios for the next three years

#### 🟢 Optimistic scenario (30%)

Zero-shot generalisation holds across the 2026 cell lines; virtual-cell screening becomes standard target-validation practice; pharma starts licensing the underlying datasets.  
  
**Consequences:** discovery timelines compress, early validation cost falls, and the nonprofit becomes the default infrastructure layer of the field. 

#### 🟡 Base scenario (50%)

Models keep improving within known tissue types, but generalisation to patient context stays partial; the Atlas keeps growing; virtual cells become one input among several in target triage.  
  
**Consequences:** incremental efficiency rather than structural change; incumbents integrate the tools without rewriting their pipeline economics. 

#### 🔴 Pessimistic scenario (20%)

Performance stalls on the next zero-shot iterations; funding flows back to wet-lab-first startups; the field consolidates around private datasets.  
  
**Consequences:** the open-benchmark advantage erodes, and the moat shrinks to the Atlas alone. 

[ Arc Institute's first virtual cell model: State Release details, training data scale, and benchmark framing from the source of the claims in this piece. Arc Institute ](https://arcinstitute.org/news/virtual-cell-model-state?ref=nexi.fund) 

Primary source — the model itself is the starting point of the analysis.

[ Arc Virtual Cell Atlas launches with data from over 300 million cells The data pipeline that underpins the virtual cell initiative, including scBaseCount and Tahoe-100M. Arc Institute ](https://arcinstitute.org/news/arc-virtual-cell-atlas-launch?ref=nexi.fund) 

The pipeline, not the model, is the durable asset.

[ 2026 Virtual Cell Challenge: zero-shot prediction in unseen cell contexts Task design, sponsors, and prize for the current benchmark round that tests generalisation honestly. Arc Institute ](https://arcinstitute.org/news/virtual-cell-challenge-2026?ref=nexi.fund) 

The 2026 rules define what "generalisation" means this cycle.

[ Arc Institute joins The Audacious Project's 2025 cohort Funding commitment behind the Virtual Cell Initiative and the billion-cell experiment target. Arc Institute ](https://arcinstitute.org/news/the-audacious-project-2025?ref=nexi.fund) 

Where the long-term capital behind the data target comes from.

[ Arc Institute launches the inaugural Virtual Cell Challenge The 2025 competition rules and the Cell commentary that started the benchmark series. Arc Institute ](https://arcinstitute.org/news/virtual-cell-challenge-2025?ref=nexi.fund) 

History of the yardstick — 2025 defines the baseline, 2026 the test.