FDA Publishes Discussion Paper on Regulation of Generative AI-Enabled Medical Devices

The U.S. Food and Drug Administration has published a discussion paper titled “Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback.”

The document was issued in August 2026 and is intended for discussion purposes only. FDA clarifies that it does not represent draft or final guidance and is not intended to propose or implement policy changes regarding how CDRH regulates generative AI-enabled devices.

Purpose of the Discussion Paper

FDA explains that generative AI-enabled medical devices may offer transformative promise for patient care and the broader health ecosystem, while also introducing unique risks when compared with traditional software and AI-enabled medical devices.

The paper is intended to seek early stakeholder input and support broader discussion on topics such as risk assessment, premarket evaluation, postmarket monitoring and other issues relevant to the regulation of GenAI-enabled devices.

GenAI-Enabled Devices and Key Challenges

FDA notes that GenAI-enabled devices may accept open-ended inputs, perform multiple subtasks and produce variable outputs to similar inputs.

They may also evolve over time through changes to models, prompts, retrieval strategies, guardrails, orchestration logic or user interfaces. Many may be built on third-party foundation models, which can offer limited transparency into training data, architecture and evaluation methods.

The paper identifies potential risks such as hallucinations or confabulations, uncertainty around intended use boundaries, limited visibility into third-party models and performance degradation across test and real-world environments.

Definitions and Regulatory Context

The paper defines GenAI as a class of AI models that emulate the structure and characteristics of input data to generate synthetic content, including images, videos, audio, text and other digital content.

FDA also discusses foundation models, large language models, multimodal systems and agentic AI systems. In this context, a GenAI-enabled device is described as a product containing one or more device software functions enabled by GenAI.

FDA emphasizes that it does not regulate GenAI as such. Instead, it regulates medical devices, including GenAI-enabled devices, using risk-based oversight where products meet the device definition under the FD&C Act.

Risk Assessment Framework

FDA presents a possible two-axis framework for thinking about risk for GenAI-enabled software functions.

The framework considers:

  • the activity performed by the device function, from non-directive informational outputs to action-taking and fully autonomous functions;

  • the consequences of relying on an incorrect device output.

According to the figure on page 7, risk increases from the lower-left to the upper-right of the framework as device activity becomes more independent and consequences become more severe.

FDA also seeks feedback on additional risk considerations, including action-directing outputs, action-taking functions, measurement and signal processing functions, patient-facing versus healthcare-professional-facing outputs, specialist versus generalist use, multi-turn conversations and care escalation functions.

Competency-Based Premarket Evaluation

FDA states that evaluation approaches developed for software with bounded inputs and fixed outputs may not be appropriate for GenAI-enabled devices.

The paper discusses a possible competency-based approach to premarket evaluation, consisting of device benchmarking followed by clinical confirmation. Under this approach, the final user-facing device, as configured for real-world deployment, would be evaluated rather than the foundation model alone.

FDA suggests that this approach could be tailored to the device’s intended use and proportionate to its risk profile.

Device Benchmarking

FDA describes device benchmarking as a scalable, high-throughput approach to non-clinical evaluation that could assess whether a GenAI-enabled device demonstrates the clinical knowledge, analytic capabilities, safety behavior, communication and generalizability needed to support reasonable assurance of safety and effectiveness.

The paper outlines potential benchmarking elements across four areas:

  • Safety — safety-critical recognition and escalation, scope maintenance, boundary adherence, calibration and uncertainty communication;

  • Clinical Proficiency — clinical knowledge, information gathering, clinical analysis, quantitative analysis and communication quality;

  • Generalizability — robustness, reliability, reproducibility and subgroup performance;

  • Agentic AI Capabilities — additional competencies for agentic GenAI-enabled devices.

The diagram on page 14 visually presents these benchmarking categories and their related elements.

Clinical Confirmation

FDA notes that benchmarking alone may not fully establish how a GenAI-enabled device will perform in real clinical use.

The paper therefore discusses clinical confirmation as a possible additional step to determine whether the device performs as intended for its intended population under real-world or clinically representative conditions. FDA also notes that clinical confirmation may not require a prospective clinical study in every case.

Potential approaches include retrospective evaluation on real patient inputs, shadow deployment, standardized patient interactions, clinician adjudication of real cases and prospective clinical studies.

Postmarket Monitoring and Change Control

FDA highlights that GenAI-enabled devices may produce varied open-ended outputs and may undergo continuous adjustments after deployment, making postmarket monitoring especially important.

Possible postmarket monitoring approaches discussed include periodic device benchmarking, sample-based clinician review and performance degradation monitoring.

The paper also discusses postmarket modifications, including sponsor-initiated changes, model evolution and changes to third-party foundation models. FDA notes that a Predetermined Change Control Plan could be one mechanism to facilitate certain device changes without requiring a new premarket submission.

Foundation Models and Agentic AI

FDA also seeks feedback on voluntary Foundation Model Device Master Files, which could allow foundation model developers and platform providers to submit information such as structured model cards or system cards to FDA.

The paper also highlights additional considerations for agentic AI systems, which can autonomously plan and execute multi-step tasks, use external tools or take actions across a sequence of steps.

Impact on Medical Device Manufacturers

For manufacturers developing GenAI-enabled medical devices, the FDA discussion paper is highly relevant because it signals the types of questions CDRH is exploring for future regulatory approaches.

Manufacturers should pay particular attention to:

  • intended use definition and scope boundaries;

  • risk assessment based on device activity and consequences of incorrect output;

  • action-directing and action-taking functions;

  • patient-facing versus HCP-facing use;

  • multi-turn conversational behavior;

  • benchmarking strategy and acceptance criteria;

  • clinical confirmation methods;

  • subgroup performance and generalizability;

  • postmarket monitoring plans;

  • change control and PCCP considerations;

  • third-party foundation model governance;

  • agentic AI safety and oversight.

For manufacturers, the key message is that GenAI-enabled medical devices may require strong lifecycle planning, robust evidence generation, transparent risk controls and continuous monitoring to support safety and effectiveness across the total product lifecycle.

Anterior
Anterior

European Commission Launches Survey on Electronic Instructions for Use for Medical Devices

Próximo
Próximo

FDA Issues Draft Guidance on Potency Assessment of Active Immunotherapy Products