In 2024, Unlearn.ai raised $50 million from Insight Partners to do something no clinical trial software company had done before: replace living, breathing patients in the placebo arm with AI-generated forecasts of how those patients would have responded had they received no treatment. The European Medicines Agency had already qualified their methodology. The FDA had given positive feedback. And Merck KGaA had signed a multi-year collaboration to use these digital twins in real immunology trials.

The placebo-controlled randomized trial has been the gold standard of clinical evidence for seventy years. It may not survive another ten.

🎯
AI-generated synthetic control arms can reduce placebo arm size by 30-50% without compromising statistical validity. That cuts trial costs, shortens timelines, and exposes fewer patients to ineffective treatments.

Regulatory acceptance is accelerating: the EMA has qualified its PROCOVA methodology for Phase 2 and 3 trials, and the FDA has accepted synthetic control arms in multiple regulatory submissions including a precedent-setting Phase 3 glioblastoma trial.

The market leaders (Unlearn, Medidata/Dassault Systèmes, and Phesi) collectively hold access to data from over 38,000 clinical trials and 12 million patients, giving them a structural advantage that makes it unlikely this remains a fragmented market.

The clinical trial cost crisis that created the opening

Bringing a single drug to market now costs over $2 billion. Roughly half of that goes to clinical development, and the single most expensive component is patient recruitment. Every patient randomized to a placebo arm represents not just an ethical cost (a person receiving no therapeutic benefit) but a logistical and financial one: screening, enrollment, monitoring, data collection, all for a subject whose data will only confirm that the experimental drug's effect is real, not whether the drug works.

In rare diseases, this model breaks entirely. Patient populations are measured in hundreds or thousands globally. A traditional two-arm randomized trial can require 60-70% of the addressable patient pool just for the control group, making the trial either infeasible or so slow that the disease progresses faster than enrollment.

The synthetic control arm solves both problems simultaneously. By generating a virtual comparator from historical trial data, real-world evidence, and AI-powered patient modeling, sponsors can reduce or even eliminate the concurrent placebo arm. The ethics improve: more patients get the experimental treatment. The economics improve: fewer sites, fewer monitors, faster timelines. And the science, its proponents argue, holds up under regulatory scrutiny.

How AI synthetic control arms actually work

The core concept is deceptively simple. Instead of recruiting and randomizing patients to receive placebo, a trial sponsor uses historical data: previous clinical trials, real-world evidence from electronic health records, and registry data, all used to construct a statistical match for the treated cohort. Each patient in the treatment arm is paired with a synthetic counterpart whose expected outcome under the control condition is predicted by an AI model trained on that historical data.

Unlearn calls this a TwinRCT, a randomized controlled trial augmented with digital twins. Their proprietary PROCOVA (prognostic covariate adjustment) methodology builds a personalized forecast for each enrolled patient: given this person's baseline characteristics, what would their disease trajectory look like without the experimental drug? The difference between the forecast and the observed outcome becomes the treatment effect estimate, measured with higher statistical precision than a traditional unadjusted analysis.

The EMA qualified PROCOVA for use in Phase 2 and 3 trials with continuous outcomes, a regulatory first for an AI methodology in clinical trial design. The FDA has provided positive feedback on the approach, and the company has published peer-reviewed validations in Nature Scientific Reports and other journals.

Its competing approach, the Synthetic Control Arm, uses a different data advantage. As a Dassault Systèmes company, it has access to patient-level data from over 38,000 clinical trials and 12 million patients, the largest such repository in the industry. Instead of building patient-level digital twins, it matches treated patients against a pool of historical control patients selected for comparable baseline demographics and disease characteristics. The FDA accepted this approach for a Phase 3 registrational trial in recurrent glioblastoma, a precedent-setting decision, and it helped the sponsor, Medicenna, reduce its enrollment requirement by two-thirds.

Phesi, the London-based AI clinical trial specialist, takes yet another path. Its Trial Accelerator platform has created digital patient profiles for over 100 million patients across 4,000 indications. These profiles serve as the foundation for digital twins that simulate patient responses in silico, enabling sponsors to test trial designs computationally before enrolling a single subject. It was recognized as a leader on Frost & Sullivan's Frost Radar for AI-enabled clinical trials in early 2026.

📊
Key signals to track

FDA final guidance on externally controlled trials. The February 2023 draft is being updated; the final version will define the regulatory ceiling for synthetic control arms.
Unlearn's next financing or acquisition. The company has raised ~$70M to date; at current pharma adoption velocity, a Series C within 12-18 months is plausible.
The first successful FDA approval based primarily on a synthetic control arm. No drug has yet crossed this line.
Phesi's IPOs or major partnership. The company's data moat (100M+ patient profiles) makes it an acquisition target for CROs.

The money behind the algorithm

The $50 million Series B in October 2025, led by Insight Partners with participation from Radical Ventures, 8VC, DCVC, and Mubadala Capital Ventures, was not an isolated event. The round reflected a growing conviction among institutional investors that AI in clinical development has moved past the pilot phase and into commercial deployment.

Its SCA is not a startup but a product line within a publicly traded parent (Dassault Systèmes traded on Euronext Paris: DSY). Its existence as an internal offering rather than a venture-backed startup signals something different: that the largest clinical technology platform sees synthetic control arms as a core service, not an experiment.

The business model is straightforward. For a sponsor running a single-arm Phase 2 trial in a rare disease, recruiting 40-60 patients instead of 80-120 saves $5-15 million in site costs alone. The synthetic control arm is priced as a fraction of that saving. The ROI is immediate and calculable, which is why top-10 pharma companies, including Merck KGaA, have already signed multi-year deals.

What limits this technology and what critics get right

Synthetic control arms are not a panacea, and their limitations are real. The statistical validity of any synthetic comparator depends entirely on the quality and relevance of the historical data used to construct it. If the historical control population differs systematically from the treated cohort on unmeasured confounders, the treatment effect estimate will be biased. Regulators therefore evaluate each application case by case, and the FDA's 2023 draft guidance on externally controlled trials makes clear that synthetic controls are not a substitute for randomization when randomization is feasible.

The EMA's qualification of its PROCOVA is instructive: it applies to Phase 2 and 3 trials with continuous outcomes. Binary endpoints (survival vs. death, remission vs. relapse) are harder to model with the same precision. And the EMA explicitly stated that qualification did not constitute approval of any specific trial design. Each use still requires regulatory review.

Data access is the other bottleneck. Medidata has the advantage of owning the world's largest repository of clinical trial patient data, an asset no startup can replicate quickly. Unlearn and Phesi have built their models on partnerships and public datasets, but the structural asymmetry means its product will always have the richest training data, while the independent players must compete on methodology rather than data scale.

💡
Who wins depends on the bottleneck that matters more

If regulatory trust is the binding constraint, Unlearn's EMA-qualified methodology and peer-reviewed validations give it the strongest moat. If data scale wins, Medidata's 38K-trial repository is irreplicable. If trial design optimization is the buyer's real need, Phesi's 100M+ patient profiles and protocol simulation capability make it the dark horse. The three approaches are complementary, not competing, and the largest CROs may acquire all three capabilities rather than bet on one.

The early-stage pipeline: startups building the next generation

Beyond the three established platforms, a wave of early-stage companies is pushing synthetic control arm technology into new territory. CellType, a Yale University spinout backed by Y Combinator's Winter 2026 batch, builds foundation models of human biology. Their Cell2Sentence model has 27 billion parameters, co-developed with Google DeepMind researchers. The company positions itself as an "agentic drug company," where AI agents autonomously simulate clinical outcomes across the entire discovery pipeline, from target identification through early clinical strategy. If their approach works at scale, it could make synthetic control arms a byproduct of a much larger simulation infrastructure rather than a standalone product.

HopeAI, based in Princeton, New Jersey, takes a different technical path. Its SynthIPD platform generates synthetic individual patient data that mimics the statistical properties of real clinical trial data. The company targets a specific pain point: many rare disease trials lack sufficient historical data to build reliable synthetic control arms using traditional matching methods. SynthIPD augments sparse real datasets with AI-generated synthetic patients, effectively creating a larger and more statistically powerful control pool than what exists in any single registry. The approach is reminiscent of how computer vision models train on synthetic images when real images are scarce.

These early-stage efforts matter because they reveal where the market is heading. The first generation of synthetic control arms solved the problem of reducing control arm size in trials with adequate historical data. That generation includes Unlearn's TwinRCT, Medidata's SCA, and Phesi's digital patient profiles. The second generation, represented by CellType and HopeAI, is trying to solve the harder problem: generating reliable synthetic controls when the historical data barely exists. That is precisely the condition that defines most rare disease trials.

The regulatory roadmap: what needs to happen next

The EMA's qualification of its PROCOVA and the FDA's acceptance of its SCA for the Medicenna glioblastoma trial are important milestones, but they are case-specific decisions, not blanket approvals. The regulatory path to mainstream adoption runs through three gates.

First, the FDA must finalize its draft guidance on externally controlled trials, originally published in February 2023. The draft laid out statistical standards for using external control arms but left significant room for interpretation on data quality requirements, patient-matching methodology, and the level of evidence needed to justify replacing a randomized control. A final version would give sponsors a predictable framework for regulatory submissions, which would in turn accelerate investment in the technology.

Second, the first drug approval based primarily on a synthetic control arm needs to happen. No regulatory submission has yet crossed this line. Every acceptance to date has involved a hybrid design where the synthetic control arm supplemented rather than replaced a randomized control. A pure synthetic-arm approval would be the signal the industry is waiting for.

Third, the data infrastructure needs to standardize. Medidata's 38K-trial repository is proprietary and not accessible to competitors. Phesi's 100M patient profiles are proprietary. The historical data that powers synthetic control arms is today the source of competitive advantage for the companies that own it, but regulators and sponsors alike would benefit from agreed-upon benchmarks for data quality, patient-matching methodology, and outcome prediction accuracy, independent of any single vendor's dataset. The FDA-EMA joint guiding principles on AI in drug development, published in January 2026, are a step in this direction, but they address AI broadly rather than synthetic control arms specifically.

What happens next

The trajectory is visible. The companies building synthetic control arms are not selling a hypothetical future capability. They are selling something pharma buyers already understand: cheaper, faster trials with the same regulatory confidence. The adoption curve in 2026 resembles where cloud infrastructure stood in 2017. The technology works, the regulatory framework is taking shape, and the early adopters have already demonstrated a return on investment that makes it harder for laggards to justify staying with the old model.

In 2022, roughly 20 AI-designed drug candidates were in clinical trials globally. By Q1 2026, that number exceeded 150, a 7x increase in four years. Every one of those trials is a potential customer for synthetic control arm technology. As the number of AI-discovered programs grows, the pressure to match discovery speed with trial efficiency will become intense.

The Lilly-Nvidia $1 billion AI Co-Innovation Lab announced at JPM 2026 is a signal in the same direction: the largest pharma companies are betting that AI can compress the entire drug development timeline, not just the discovery phase. Synthetic control arms are the most mature application of that thesis on the clinical side, and the one closest to widespread adoption.

As we wrote in July, the FDA is ready to listen on synthetic evidence for rare diseases. What has changed since then is that the technology is no longer waiting for permission. It is already being used in active regulatory submissions, and the evidence base for its validity is growing faster than the guidance that governs it. The question is no longer whether synthetic control arms will become standard in rare disease trials. The question is which approach will define the standard: Medidata's data scale, Unlearn's methodological rigor, or Phesi's simulation capability. The most likely answer is that the market does not converge on one. It fragments into use-case-specific solutions, which would be the healthiest outcome for sponsors but the hardest for regulators to oversee.

Clinical trials gain intelligence
Nature Biotechnology survey of AI in clinical development — covers synthetic control arms, digital twins, and regulatory acceptance. The key source for understanding how the field evolved from 2022 to 2025.
Comprehensive overview of how AI-generated synthetic controls and digital twins are reshaping clinical trial design
Synthetic Control Arm in Clinical Trials
Medidata's Synthetic Control Arm product page with detailed FAQ on FDA acceptance, data sources (38K+ trials, 12M+ patients), and the Medicenna rGBM precedent.
The most detailed public documentation of a commercially deployed synthetic control arm platform with real regulatory precedent
Digital Twins in Clinical Trials: Virtual Controls and FDA
Independent analysis of the regulatory landscape for digital twin and synthetic control arm technologies in clinical trials, including FDA and EMA guidance documents.
Independent regulatory analysis that connects individual company developments to the broader FDA/EMA guidance evolution