Field guide · 2026-07-22
FDA's AI guidance, explained for people who run trials
FDA counted more than 500 CDER submissions with AI components since 2016, and CBER counted more than 560 through 2023. The agency was reviewing AI-derived evidence for years before it wrote down how. The write-down finally arrived on January 6, 2025: a draft guidance called Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products, the first FDA guidance on AI in drug and biologic development.
Eighteen months later it is still a draft. The comment period closed April 7, 2025, and as of mid-2026 no final version has been published. That has not stopped sponsors from treating it as the operative playbook, because it describes how FDA reviewers already think. If you work on a trial where a model's output feeds a submission, this document is the closest thing to the rules.
Here is what it actually says, without the law-firm alert padding.
The one question that decides whether you care
The guidance covers AI models that produce information supporting a regulatory decision about a drug's safety, effectiveness, or quality. That reaches across the lifecycle: predicting patient outcomes, disease-progression models, dose selection, manufacturing process controls, pharmacovigilance analytics, trial-design support.
It explicitly excludes two things. Drug discovery (target identification, lead optimization, molecule screening) is out. So are operational efficiencies, the Federal Register notice's term for internal uses that do not affect patient safety, drug quality, or the reliability of study results.
The practical test: if your AI touches a claim about safety, efficacy, or quality in a submission, you are in scope. If it just makes your team faster, you are out. A model that drafts your monitoring visit letters is FDA's business only in the sense that nothing is; a model that imputes missing endpoint data is squarely covered.
The two-factor risk test
Everything in the framework scales to two definitions.
Context of use is the model's specific role and scope: what it will and will not do, and how its output feeds the regulatory question alongside other evidence. The guidance's whole architecture hangs on writing this down precisely before anything else.
Model risk is then a function of two dimensions. Model influence measures how much weight the model's output carries in the decision relative to other evidence. Decision consequence measures what happens if a decision informed by a wrong output goes wrong: patient safety, product quality, reliability of results. High influence combined with high consequence means high model risk, which means more credibility evidence. A model that is one input among many into a low-stakes call needs little. A model that alone drives a dosing decision needs a lot.
This is the piece worth internalizing, because it is how a reviewer will size up any AI in your submission. Not "is it AI" but "how much is riding on it, and how bad is a miss."
The seven steps
The credibility assessment framework runs:
- Define the question of interest, meaning the specific development decision the output will inform.
- Define the context of use.
- Assess the model risk, from influence and consequence.
- Develop a credibility assessment plan: data, training, validation approach, acceptance criteria, all proportionate to the risk.
- Execute the plan.
- Document the results in a credibility assessment report, including deviations.
- Determine whether the model is adequate for its context of use. If not: gather more evidence, reduce the model's influence on the decision, or narrow the context of use.
Step 7 is quietly the most useful. It gives you three exits when validation comes up short, and two of them do not require a better model. You can demote the model to one input among several, or shrink what you claim it does.
What you actually have to produce
A credibility assessment is not a certificate. It is a documented evidence package, and for an in-scope model it looks like this in practice: a written statement of the question of interest and context of use, a model-risk determination, a plan written before execution that pins down training and test data provenance (FDA's phrase is "fit for use"), performance metrics with pre-specified acceptance criteria, the validation results, the report, and a description of human oversight.
FDA also signals, repeatedly, that it wants to see sponsors early: INTERACT, pre-IND, and milestone meetings, with the credibility plan agreed before the evidence is built rather than argued about after. And credibility is not one-and-done. The guidance expects lifecycle maintenance, meaning monitoring for performance change after deployment.
None of this floats free of the rules you already know. 21 CFR Part 11 still governs the electronic-records layer of any AI tool that generates or holds regulated records: audit trails, access controls, system validation. ICH E6(R3), which FDA adopted in 2025, contains no AI-specific language, but its principles already apply: sponsor accountability that does not transfer to a vendor, risk-proportionate quality management, end-to-end data traceability. For AI-assisted work that means records showing the source, human review, and acceptance of anything a model wrote into the trial record. Part 11 and E6(R3) do not mention AI, but they already govern it. The credibility framework is a new layer on top, not a replacement. If E6(R3) is unfamiliar territory, the risk-based monitoring shift is covered separately.
The sponsor point deserves its own sentence: the oversight burden is non-delegable. When the CRO runs the tool, or a software vendor supplies the model, the credibility problem still belongs to the sponsor.
What has happened since January 2025
Two developments matter.
In April 2026, FDA published a request for information on an "AI-Enabled Optimization of Early-Phase Clinical Trials" pilot program. The agency calls early-phase trials a "critical bottleneck" and asks industry how AI could improve patient selection, dose optimization, safety monitoring, and go/no-go timing, and how a pilot should be scoped, run, and measured. Comments closed June 29, 2026 after an extension. Read the posture shift: the 2025 guidance is FDA deciding how to judge your AI; the 2026 RFI is FDA asking how AI should reshape trials it oversees.
Meanwhile FDA deployed its own. Elsa, the agency's internal generative-AI tool, launched June 2, 2025 and rolled out agency-wide within the month, running inside a government cloud, not training on sponsor data, and used for scientific review support, protocol assessment, and inspection targeting, with further upgrades reported in 2026. The agency writing the credibility rules is itself using AI to review your submissions. Whatever you conclude from that, it is not an agency hostile to the technology.
Three documents people confuse with this one
The device guidance. The same week, FDA's device center issued a separate draft on AI-enabled device software functions. That track regulates AI as the product: total-lifecycle management, bias analysis, and predetermined change control plans that pre-authorize model updates. The drug guidance regulates AI as evidence supporting a product's claim. Different centers, different logic. If a vendor cites clearance under the device framework as proof their tool satisfies the drug framework, that is a category error.
The 2023 discussion papers. FDA published two in May 2023, one on AI in drug development and one on manufacturing. They were consultation documents, not guidance; the 800-plus comments they drew shaped the 2025 framework. If someone cites "FDA's 2023 AI guidance," this is what they are misremembering.
The JAMA viewpoint. In October 2024, FDA officials including then-Commissioner Robert Califf published a JAMA piece arguing that the scale of effort needed to continuously re-evaluate AI models "could be beyond any current regulatory scheme." That is officials thinking out loud about whether the traditional model scales, not policy. It is also the frankest sentence anyone at the agency has produced on the subject.
How Europe differs
EMA got there first and finished. Its reflection paper on AI across the medicinal product lifecycle was adopted September 9, 2024, final rather than draft. It is principles-based: risk-proportionate expectations, human-centric design, data quality, transparency, early engagement for higher-risk uses. It also spans the full lifecycle including discovery, which FDA excluded.
The structural difference: FDA's draft hands you an operational seven-step procedure; EMA's final paper hands you principles and promises detailed guidelines later. Layered on top of EMA sits the EU AI Act, in force since August 2024, a horizontal law that can apply to the same tool in parallel with medicines regulation. The US has no equivalent statute, so FDA works entirely inside its existing authorities. A sponsor running a global trial with one AI tool can face the FDA framework, EMA expectations, and the AI Act at once.
The questions teams are actually asking
Across the practitioner and law-firm analyses, the same questions recur. Is my use in scope, evidence or efficiency? How much validation is enough for my risk tier? What counts as adequate provenance for a model trained on a vendor's data? Who owns credibility when the CRO supplies the model? Do I need Part 11 validation and a credibility report? When do I have to talk to FDA before submitting?
The guidance answers the first and last cleanly: the scope line above, and "earlier than you think" via the engagement pathways. The middle four all resolve to the same move. Write the context of use narrowly, score the model risk honestly, and let the answer scale from there.
One caution for anything you read on this topic, including vendor summaries: check the status date. A final guidance could publish any month and change details. As of July 2026, the draft is what exists, and the two-factor risk test at its center is not the part likely to change. For the wider picture of what AI is actually doing in trials, start with AI in clinical research: what's real.
Discussion
0 commentsAnonymous, verified members. House rules apply.Nobody has weighed in yet. If this piece matches or misses your experience, say so below; one sentence is enough.
Open Label is building the salary dataset this industry never had. Add your anonymous datapoint. Three minutes, no name, no email.