July 28, 2026
AI is showing up in the part of a clinical trial that draws the most regulatory scrutiny: the analysis itself, where a model's prediction becomes part of the treatment effect estimate. It's a shift Unlearn helped set in motion. For the teams designing those trials, it's critical to understand which use cases will hold up when a regulator asks them to defend the design.
That question has a clearer answer than most sponsors expect. And the place to start is a framework the FDA has already published.
A common framework for an uncommon problem
In the FDA’s guidance document, Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products, the agency lays out a seven-step process for assessing whether a model is credible enough for the job it's being asked to do. The guidance was written for safety, efficacy, and quality decisions, but the logic travels well beyond that scope. It asks you to define the specific question the model is answering, the exact context in which it's used, and then to weigh risk along two axes: how much the model influences the answer, and how serious the consequences are if that answer is wrong.
That second pairing is what makes the framework useful. A model can exert significant influence on a result and still be low risk if being wrong costs little. Another can play a smaller role and still demand rigorous validation, because the cost of error hurts the credibility of a registrational trial. Risk isn't a property of "using AI." It's a property of a specific use, in a specific trial, answering a specific question.
Three uses, three different risk profiles
In a recent whitepaper, our team applies this framework to three uses of AI in trial analyses that we've worked through in real trials. Each one puts a model's predictions directly into the treatment effect estimate, and each lands in a very different place on the risk spectrum.
Some sit on firm ground. Adding power to an RCT through prognostic covariate adjustment produces a valid treatment effect estimate whether or not the model is prognostic; a useful model tightens the estimate, and a weak one costs you almost nothing. It's a method the EMA has qualified and that the FDA has supported. Using that same approach to prospectively reduce sample size raises the stakes a step, because now the design depends on the power the model is expected to deliver, which makes validation against relevant data essential.
Others carry risk that needs to be characterized carefully. Using model-based comparators to support a single-arm study, in which predicted control outcomes serve as a control group, offers real benefit to patients but no guarantee of unbiased estimation. That doesn't rule it out. It raises the bar for transparency and validation, and it's exactly the kind of use the credibility framework was built to interrogate before a trial begins, not after.
What you can tell before the trial starts
Our whitepaper aims to give clinical and biostatistics teams a way to assess its use against what regulators will expect to see, and to flag where the path is still being written. Model-based synthetic controls, for instance, hold a real advantage over data-based comparators, since a model can be validated prospectively, but their regulatory path is still forming.
If your team is weighing where AI fits in an upcoming trial, this is a practical place to start: which applications are defensible today, which need more evidence, and how to tell the difference before you commit to a design.
