September 30, 2026
Most people using AI today never had to learn the core principles of how machine learning works. Useful, natural AI interfaces like Claude, ChatGPT, and other large language models have become prevalent faster than understanding of the technology has spread. This gap shows up everywhere in how pharma builds, adopts, and values AI.
“Chat is all you need” is fine for a general question, but with the mountains of private data in healthcare, and especially for the deeper scientific questions that come up in drug development, I’d argue that’s not the case.
There’s no codified set of principles behind machine learning – they’re learned by doing it. If I had to write them down, there are three:
1. Build large, relevant datasets
2. Build general, large models, then fine-tune them for specific use cases
3. Validate for the use case(s) of interest
These may feel natural, but we often see teams do the opposite. There is a strong inclination to only use data narrowly tailored for use cases, or models that are simple and fit for a single purpose only, or validation that just asks about “bias” generally without thinking about how the model will be involved in the decision making process. Although the first two aren’t hard and fast rules – the third is a core tenet of FDA’s draft guidance on AI in drug development, and just makes sense – they should be the default approach, not the exception.
If you buy into these principles, it’s natural to buy into the approach: Building rich data and models creates a computational tool you can use for many applications. Invest in the datasets and the models, and you can tap them when you need them.
This is what digital twins are – a computational model that comprehensively predicts the outcomes of individual patients on a defined standard of care. We collect data relevant to patient outcomes to build a full picture of disease course. We build models that capture this picture, even when it’s complex and multi-faceted. And when we apply the digital twin model to a specific use case, we fine tune as needed and validate for that application. Principles one, two, and three.
So what can we do with digital twins? Every trial is a chain of decisions,and you don't have to wait for the next protocol to start making them with better evidence. There are three decision points where a team can come in.
IN DESIGN, BEFORE FPI
In planning, the main job of the clinical development team is to arrive at a design that achieves goals for evidence generation and is also feasible, meaning it can run with a reasonable budget in a reasonable time, and can stand up to internal governance. That evidence for feasibility can be found in patient-level datasets from past clinical trials, observational studies, and relevant RWE, and in simulations from models built on this data. It lets us answer questions like how variability changes across the population, expected control arm outcomes, how endpoints relate to each other and which ones to select, and what overall power and sample size are really needed.
Some trial design choices matter very little for overall power but meaningfully improve enrollment rates. Likewise, there may be selection rules that sharpen the study’s ability to answer key questions about the mechanism of action, efficacy, and safety. When choices are backed up with data and simulations, it strengthens the case for governance and makes it much easier to get alignment across the team and with leadership.
IN PROGRESS, FPI THROUGH READOUT
Once the study is running, there’s lots to do for trial operations to ensure the quality of data. The protocol is locked, but the decisions aren't over: whether to stop for efficacy or futility, whether something in the data needs a closer look, or whether additional sensitivity or exploratory analyses are warranted. This is where models can be useful to provide context and help interpret the data as it’s coming in. For example, we can use digital twin predictions for enrolled patients to understand what the expected control arm outcome will look like, or to provide a synthetic control reference in open-label or early-phase studies with small control arms. This helps teams understand options for future trials, and digital twins can also support interim analyses for stronger signals of futility or early efficacy. For data quality, models and reference data are useful to flag overly strong placebo effects at the study, site, or rater level, and can be used to detect data quality issues and correct them weeks or months earlier.
AFTER READOUT, INTO THE NEXT STUDY
Finally, and probably most importantly, is the analysis phase of the trial. The question at readout is how much information you can get out of the data you already collected, because that decides whether the program continues. We have done a lot of work to define ways to include model predictions in trial analyses and improve statistical power, efficiency, and the robustness of results. For randomized controlled trials (RCTs), digital twins can be incorporated into analyses as covariates using prognostic covariate adjustment (PROCOVA), an approach we took through EMA qualification in 2022 and one frequently cited by FDA staff as an example of a low-risk application of AI when they speak about FDA’s draft AI guidance. PROCOVA is in wide use, and we have published on it through use cases in trials as well as recommendations for implementation under the AI guidance. There are other ways to use digital twins in the analysis of RCTs, including Bayesian analyses, but digital twins can also be used as synthetic controls in single-arm or open-label trials. All told, there is a rich set of ways to benefit studies built on the back of a solid understanding of data and models, and what you learn at readout is what you bring to the next design.
Why the principles don’t get applied
I’ll close by re-emphasizing that much of the data these applications are built on is private, or hard to access and potentially expensive. Using it takes focused effort, but the competitive advantage that digital twins bring makes the effort worth it.
Machine learning principles are new in pharma, and organizational change can be difficult. What closes that gap is experience and validation. This is where we partner heavily with sponsors. We bring experience with data, models, and digital twins. More importantly,we bring the experience to make applications that work, even regulated ones. PROCOVA is the clearest example, but the same applies to external controls, trial planning, and additional specialized applications we haven’t covered here.
The future is one where tools to use data and models are widely available and teams are using them heavily in decision-making. We increasingly partner with sponsors to not just build but enable – just like everyone can open a chat window and get answers to general questions, we want everyone to open their datasets and use digital twins to guide their decision-making in drug development.
If you see the future of clinical development as one that embraces how these principles can meaningfully improve clinical trials, let's talk about how to get there.
