Generative AI Medical Devices: What the FDA Is Asking

Generative AI medical devices face a new FDA approach. What the discussion paper proposes on risk, competency testing, and monitoring, and how to prepare.

Generative AI Medical Devices: Reading the FDA Discussion Paper

Quick Answer: The FDA's August discussion paper on generative AI medical devices proposes a two-axis risk framework, premarket competency assessment combining benchmarks with clinical confirmation, and stronger postmarket monitoring, including for foundation models and agents. Comments close October 19. Manufacturers should start building evaluation evidence now.

Generative AI medical devices are about to get a regulatory framework of their own. In August, the FDA published a discussion paper titled Considerations for the Regulation of Generative AI-Enabled Medical Devices and asked for feedback, with comments due by October 19 under docket FDA-2026-N-7874. The paper proposes a two-axis approach to assessing risk, a premarket "competency assessment" that combines non-clinical benchmarking with clinical confirmation, new thinking on postmarket monitoring, and specific attention to foundation models and agentic systems. It is not yet guidance, and nothing in it is binding. But it is the clearest signal so far of what evidence the agency will expect, and manufacturers who wait for final rules will be starting late.

Traditional AI Devices Versus Generative AI Devices

Most AI-enabled devices on the market produce narrow, fixed outputs. Generative systems do not, and that changes almost every part of the evidence picture.

FactorTraditional AI-Enabled DeviceGenerative AI-Enabled Device
OutputA score, flag, or classificationOpen-ended text, summaries, or images
Typical failureA missed or false detectionFabrication, omission, or inconsistent answers
Premarket evaluationPerformance on a fixed test setBenchmarks across tasks plus clinical confirmation
Change over timeLocked, or managed through a change control planUnderlying foundation models may be updated by third parties
User interactionAn alert or number to interpretA conversation that shapes clinical reasoning
Postmarket signalMeasurable drift in accuracyContent errors that are harder to detect at scale

What the FDA Published

The FDA's public list of authorised AI-enabled devices now runs to well over a thousand products, and reports put the figure above 1,600. The overwhelming majority are in imaging and signal analysis, producing outputs a clinician can check against a defined standard. Generative products, such as tools that draft clinical notes, summarise records, answer clinical questions, or generate patient-facing explanations, have largely sat outside that list or been positioned to avoid device classification. The agency's digital health advisory committee examined generative AI devices in 2024, and the new discussion paper builds on that work.

The paper does four things. It proposes assessing risk along two axes rather than one, recognising that a generative tool's risk depends on more than its intended clinical area. It outlines competency assessment as the premarket evidence model. It explores how postmarket monitoring should work when errors are content rather than numbers. And it asks targeted questions about foundation models and agentic AI. Manufacturers should read the definitions in the paper itself, because the framing of those axes will shape how products are positioned.

Why Generative AI Medical Devices Are Different

A classic diagnostic algorithm can be wrong, but it is wrong in a bounded way. It misses a nodule or flags one that is not there, and both can be counted against a reference standard. Generative systems are wrong in open-ended ways, and the same input can produce different outputs on different runs.

How Generative AI Medical Devices Fail

  • Fabrication. A summary includes a medication, finding, or history item that is not in the source record.
  • Omission. A critical allergy, contraindication, or abnormal result is left out of an otherwise accurate summary.
  • Inconsistency. The same question asked twice produces materially different answers.
  • Prompt sensitivity. Small changes in wording lead to large changes in output.
  • Upstream model change. An update to the foundation model alters behaviour without any change by the manufacturer.
  • Automation bias. Fluent, confident text encourages clinicians to accept it with less scrutiny than they would a junior colleague's work.

See How Consistently Models Answer

Run the same clinical research question across six AI models and see where their answers diverge.

Try Talkory Free

Competency Assessment in Plain Terms

The idea behind competency assessment is closer to how we judge a clinician than how we validate a lab test. Rather than one accuracy figure on one dataset, a manufacturer would show that the system performs acceptably across the range of tasks it is meant to do, using structured non-clinical benchmarks, and then confirm in a clinical setting that the benchmarked performance holds up with real users and real data.

In practice that suggests a different evaluation programme from the one most device teams are used to. It means defining the tasks the product performs, building test sets that cover typical and difficult cases, writing scoring rubrics that capture fabrication and omission as well as overall quality, using qualified clinical reviewers, measuring variation across repeated runs, and then running a clinical study that checks the benchmarks predict real-world performance. None of this is trivial, and much of it is expensive to do retrospectively.

Preparing Now: Seven Steps

  1. Write a precise intended use. Which tasks, which users, which settings, and what the output is used for.
  2. Map your risk on both axes. Use the paper's framing to understand where your product is likely to sit.
  3. Build a task-based benchmark. Cover routine and edge cases, with rubrics for fabrication, omission, and safety-relevant errors.
  4. Measure run-to-run variation. Report how much outputs vary for the same input, not just average quality.
  5. Plan for foundation model change. Decide how you will detect, assess, and control updates you do not make yourself.
  6. Design postmarket monitoring for content errors. Sampling, clinician feedback channels, and defined error categories.
  7. Submit comments or engage early. The comment window and pre-submission meetings are opportunities to shape expectations.

Pros and Cons of the Proposed Approach

  • Pro: a clearer path to market. A defined evidence model reduces uncertainty for developers who want to do this properly.
  • Pro: evaluation that fits the technology. Task-based competency testing captures failures a single accuracy number would miss.
  • Pro: attention to foundation models. Third-party model dependency is a real risk that deserves explicit treatment.
  • Con: higher evidence costs. Benchmarks, expert reviewers, and clinical confirmation add time and expense.
  • Con: moving definitions. Until guidance is final, manufacturers must plan against a framework that may change.
  • Con: monitoring is hard. Detecting content errors after deployment is far harder than tracking a performance metric.

Real Scenarios Worth Thinking Through

These scenarios are illustrative, showing how generative AI medical devices play out in practice rather than presented as verified case studies.

A company building a discharge summary tool reports high overall quality scores from clinician reviewers. When it adds a rubric for omissions, it finds that a small share of summaries leave out a medication change. The average score hid the safety-relevant failure, which is exactly what task-based competency testing is designed to expose.

A manufacturer relies on a third-party foundation model. The provider releases an update that improves general fluency but subtly changes how lab values are reported. Because the manufacturer runs its benchmark on every model version, the change is caught before it reaches clinicians.

A start-up positions its clinical question-answering tool as non-device decision support. After reading the discussion paper, its regulatory lead concludes that the way clinicians actually use it may not support that position, and the company begins building device-grade evidence early rather than late.

Evaluating Models on Sensitive Data?

Enterprise plans cover private deployment, custom data residency, dedicated infrastructure, and an SLA.

Talk to Enterprise Sales
“After testing multiple AI models on coding, research, and business prompts, combined outputs produced more reliable results than any single model.” Internal multi-model evaluation, Talkory research team.

Foundation Models, Agents, and the Wider Picture

Two questions in the paper deserve particular attention. The first is foundation models. Most generative devices will be built on models the manufacturer did not train and cannot fully inspect, which shifts the focus to supplier agreements, version control, and evaluation of every update. The second is agentic AI: systems that take actions, not just produce text. An agent that orders a test or updates a record raises risk in a way a drafting tool does not.

Manufacturers selling in Europe face a parallel track, since AI in regulated medical devices falls within the EU's high-risk category, with timing still in flux as we covered in EU AI Act compliance after the omnibus delay. For AI used in development rather than in the device itself, the expectations look different again, covered in AI in clinical trials and validating GenAI in pharma.

Why Talkory Wins

Talkory is not a medical device and is not intended for clinical decisions. Where it helps is the research and evaluation work that sits behind a regulatory strategy. It runs the same question across GPT, Claude, Gemini, Grok, Perplexity Sonar, and Kimi K3 in one pass, which makes inconsistency between models immediately visible. Regulatory and evaluation teams can use that to test how well different models handle a task before choosing a foundation model, to find the questions where model answers diverge most and therefore belong in a benchmark, and to cross-check their reading of a regulatory question before taking it to counsel or the agency.

Final Verdict

Generative AI medical devices are moving from a regulatory grey zone toward a defined evidence model. The FDA's discussion paper points to two-axis risk assessment, competency assessment that pairs benchmarks with clinical confirmation, monitoring designed for content errors, and explicit controls for foundation models and agents. The details will change before anything is final. The direction is clear enough to act on: write a precise intended use, build task-based benchmarks with rubrics for fabrication and omission, measure variation, plan for upstream model change, and engage with the agency while the framework is still being shaped.

Find Where Models Disagree

Compare six AI models on the same question and see which answers belong in your benchmark.

Try Talkory Free

Frequently Asked Questions

What did the FDA publish on generative AI medical devices?

In August the FDA released a discussion paper, Considerations for the Regulation of Generative AI-Enabled Medical Devices, requesting public feedback. It covers risk assessment, premarket competency assessment, postmarket monitoring, foundation models, and agentic AI. Comments are due by October 19.

Is the FDA discussion paper binding guidance?

No. It is a discussion paper with a request for feedback, not final or draft guidance. It signals the agency's current thinking and the kind of evidence it is likely to expect, so manufacturers should plan with it in mind.

What is competency assessment for AI devices?

It is the proposed premarket evidence model combining non-clinical benchmarking across the tasks a system performs with clinical confirmation that performance holds in real use. It resembles assessing a clinician's competence more than validating a single lab test.

How do generative AI medical devices fail differently?

They can fabricate information, omit critical details, give inconsistent answers to the same input, react strongly to wording changes, and change behaviour when an underlying foundation model is updated. These failures are harder to measure than classification errors.

How should manufacturers prepare now?

Define a precise intended use, map risk using the paper's framing, build task-based benchmarks with rubrics for fabrication and omission, measure run-to-run variation, plan controls for foundation model updates, and design postmarket monitoring for content errors.

MB

Mital Bhayani, AI Researcher & SaaS Growth Specialist

Mital writes on multi-model AI accuracy, SaaS growth, and AI governance in scientific research and life sciences. Reviewed by Chetan Kajavadra, Lead AI Researcher at Talkory.ai. Connect on LinkedIn →

๐Ÿค–

Get 5 AI perspectives on this topic

Talkory runs your question through GPT, Claude, Gemini, Grok, Sonar & Kimi K3 simultaneously, then cross-checks the answers.

Try Talkory.ai free โ†’
โ† Back to all articles

Related Articles

๐ŸงฌLife Sciences

AI in Clinical Trials: What Inspectors Will Ask

The question has moved from whether AI is allowed in trials to how a sponsor explains its use when an inspector asks, and that conversation now happens long before the marketing application.

Read article โ†’
๐Ÿ“ฐAI and Media

Can AI Spot Fake News? We Tested All 5 Models

We built a 20-headline test, half real and half fake, and ran it through ChatGPT, Claude, Gemini, Grok, and Perplexity. Claude scored 90%. Grok scored 70% while sounding 95% confident. Confidence without accuracy is the failure mode that actually spreads misinformation.

Read article โ†’
โœˆ๏ธAI Travel

Best AI for Travel Planning: We Tested All 5 Models

We gave all five AI models the same Tokyo prompt and audited every restaurant, museum, and transit direction. Perplexity scored 95%. Grok scored 63%. A hallucinated restaurant ruins a vacation. Here is what the field looks like.

Read article โ†’
๐Ÿ’ฐAI for Finance

We Asked 5 AI Models to Build a $10K Portfolio

Five models. Same prompt. One $10,000 portfolio test. Gemini returned the most. Claude managed risk the best. Perplexity was the easiest to defend. And the disagreements between them told us more than any single answer could.

Read article โ†’
๐Ÿค–

Stop guessing. Get verified AI answers.

Talkory.ai queries GPT, Claude, Gemini, Grok, Sonar and Kimi K3 simultaneously, cross-verifies their answers, and gives you a confidence-scored consensus. Free to start.

โœ“ Free plan includedโœ“ No credit cardโœ“ Results in seconds