Regulatory Intelligence Brief · Clinical Trials & AI
FDA's Real-Time Clinical Trial Initiative and the Limits of AI-Enabled Regulatory Review
Real-time data exchange may reduce delay, but AI-supported regulatory decisions require separate, context-specific validation.
A Faster Data Path Is Only the Beginning
On April 28, 2026, the U.S. Food and Drug Administration announced the successful initiation of two proof-of-concept real-time clinical trials. FDA Commissioner Marty Makary described the conventional development model as containing long periods of “dead time” between trial phases and framed real-time data exchange as one way to reduce those delays.1
The basic idea is straightforward. Clinical trial data ordinarily move from investigative sites through sponsor systems, data cleaning, analysis, and regulatory packaging before reaching FDA. In describing the initiative, FDA asked what information reviewers actually need during a trial and proposed that defined endpoints or safety signals could be transmitted as the study progresses rather than only after the full dataset has been processed.2
Real-time signal sharing and AI-supported decision-making, however, are not the same regulatory problem. FDA's April 29 request for information proposed a broader pilot involving AI-enabled optimization of early-phase trials and identified possible applications including recruitment, dose escalation, safety monitoring, adaptive designs, earlier Phase 1-to-Phase 2 decisions, biomarker assessment, patient selection, and endpoint validation.3
AI Is Not One Regulatory Problem
These possible uses should not be treated as a single technology category. An algorithm used to identify potential trial participants raises different questions from one used to summarize adverse events. A dose-escalation model can influence immediate participant exposure, while a biomarker model can affect eligibility, stratification, and the population to which study results apply. Endpoint validation introduces still another set of evidentiary questions. Each context requires its own data controls, validation strategy, performance measures, and level of human review.
Trial design adds another layer of variation. A first-in-human oncology study, a healthy-volunteer study, and an adaptive rare-disease trial differ in sample size, baseline uncertainty, missingness, time sensitivity, and tolerance for error. FDA's request appropriately asked which trial types and development settings might benefit most from an AI-enabled approach rather than presuming that a single model or control framework would work across all early-phase research.3
Long-Context Performance Requires Testing
One practical challenge is the difficulty some large language models have when important information appears within long inputs. Research on “lost in the middle” behavior found that performance can decline when relevant material is positioned between the beginning and end of a lengthy context.4 This limitation is most relevant to language-model retrieval, summarization, and document-review use cases; it should not be generalized to every form of AI used in a clinical trial.
In real-time safety monitoring, however, the limitation matters. A meaningful risk pattern may be distributed across an investigator narrative, laboratory trends, prior adverse events, concomitant medications, protocol deviations, and data from other sites. If a language model retrieves or summarizes only part of that record, it may obscure the pattern that requires attention. Validation should therefore test realistic long-context inputs, information placement, missing data, conflicting evidence, and performance at the point where the output will actually be used.
Trustworthy AI Depends on Context of Use
The National Institute of Standards and Technology's voluntary AI Risk Management Framework describes trustworthy AI through several connected characteristics: validity and reliability, safety, security and resilience, accountability and transparency, explainability and interpretability, privacy enhancement, and fairness with harmful bias managed.5 These characteristics may pull in different directions. A recruitment model may improve efficiency while worsening representation if its training data reflect unequal access to care. A biomarker model may improve stratification while creating privacy risks or performing poorly in underrepresented subgroups.
FDA's January 2025 draft guidance on AI supports a similarly cautious view. The document is Draft Level 1 guidance and is not for implementation. It recommends a risk-based credibility assessment for an AI model in a defined context of use and explains model risk through both the model's influence on a decision and the consequences of an erroneous decision.6 The central lesson is that credibility is context-specific: evidence sufficient for an administrative support tool may be inadequate for a model that influences dosing or a regulatory conclusion.
Implementation Details Matter
FDA's internal use of generative AI illustrates both the potential and the governance challenge. The agency launched Elsa in 2025 and later introduced Elsa 4.0 with the HALO data platform, describing capabilities that include document search, protocol review support, adverse-event summarization, quantitative analysis, visualization, and access to more than 40 application and submission data sources.7 Better retrieval and workflow integration could reduce repetitive work, but efficiency does not remove the need to verify outputs.
A July 2025 trade-press report, citing anonymous FDA personnel, raised concerns about inaccurate outputs, including fabricated citations. Those allegations were not established in an official FDA finding, but they illustrate a known generative-AI risk and the importance of validation, audit trails, source traceability, and meaningful human oversight.8
Continuity also matters. Vendors, base models, cloud environments, retrieval systems, prompts, and data connectors can change during a development program. Even when the use case remains the same, a material system change can alter performance. Consistent with FDA-EMA lifecycle principles, sponsors should document the change, assess its impact, perform targeted regression testing, and determine whether proportionate revalidation is necessary.9
A Stepwise, Risk-Based Path
The proposed pilot should proceed stepwise and tie controls to the influence of the model and the consequences of error. As an illustrative—not FDA-defined—stratification, lower-risk uses might include document search, missing-field detection, and operational dashboards. Medium-risk uses might include recruitment support, protocol-deviation triage, and draft adverse-event summaries, provided that humans review the output before action. Higher-risk uses would include safety-signal prioritization, dose-escalation recommendations, biomarker-based selection, adaptive-design decisions, endpoint validation, or recommendations to move from Phase 1 to Phase 2. These uses warrant stronger prospective validation, uncertainty reporting, subgroup testing, auditability, change control, and documented human decision authority.10
Conclusion
FDA's real-time clinical trial initiative is a meaningful modernization effort. Sending defined signals to reviewers as a trial progresses may reduce avoidable delay and support earlier regulatory dialogue. But the AI portion of the proposal is not solved by faster data transmission. Any AI-enabled use must be valid for its context, sufficiently explainable, fair across relevant groups, privacy-conscious, reliable under realistic conditions, and subject to appropriate human control. The most credible path is to pilot narrowly, measure performance against clear comparators, learn from failures, and expand only when the evidence supports doing so.
References
- U.S. Food and Drug Administration. “FDA Announces Major Steps to Implement Real-Time Clinical Trials.” April 28, 2026.
- U.S. Food and Drug Administration. “FDA Direct: The Power of Real-Time Clinical Trials.” April 29, 2026.
- U.S. Food and Drug Administration. “AI-Enabled Optimization of Early-Phase Clinical Trials Pilot Program; Request for Information.” 91 Fed. Reg. 23100 (April 29, 2026).
- Liu, Nelson F., et al. “Lost in the Middle: How Language Models Use Long Contexts.” Transactions of the Association for Computational Linguistics 12 (2024): 157–173.
- National Institute of Standards and Technology. “Artificial Intelligence Risk Management Framework (AI RMF 1.0).” NIST AI 100-1, January 2023.
- U.S. Food and Drug Administration. “Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products.” Draft Guidance, January 2025.
- U.S. Food and Drug Administration. “FDA Launches Agency-Wide AI Tool to Optimize Performance for the American People.” June 2, 2025; and “FDA Expands AI Capabilities and Completes Data Platform Consolidation.” May 6, 2026.
- Applied Clinical Trials. “FDA's Elsa AI Tool Raises Accuracy and Oversight Concerns.” July 2025.
- U.S. Food and Drug Administration and European Medicines Agency. “Guiding Principles of Good AI Practice in Drug Development.” January 14, 2026.
- The use-case stratification is the author's illustrative application of the risk-based principles in references 5 and 6; it is not an FDA classification.
Independent portfolio analysis based on publicly available information. This article was not prepared for or endorsed by FDA, the University of Georgia, AstraZeneca, Amgen, or any other organization and does not constitute legal, medical, or regulatory advice.
← Back to portfolio