Generative AI Medical Devices: Reading the FDA Discussion Paper
Generative AI medical devices are about to get a regulatory framework of their own. In August, the FDA published a discussion paper titled Considerations for the Regulation of Generative AI-Enabled Medical Devices and asked for feedback, with comments due by October 19 under docket FDA-2026-N-7874. The paper proposes a two-axis approach to assessing risk, a premarket "competency assessment" that combines non-clinical benchmarking with clinical confirmation, new thinking on postmarket monitoring, and specific attention to foundation models and agentic systems. It is not yet guidance, and nothing in it is binding. But it is the clearest signal so far of what evidence the agency will expect, and manufacturers who wait for final rules will be starting late.
Traditional AI Devices Versus Generative AI Devices
Most AI-enabled devices on the market produce narrow, fixed outputs. Generative systems do not, and that changes almost every part of the evidence picture.
| Factor | Traditional AI-Enabled Device | Generative AI-Enabled Device |
|---|---|---|
| Output | A score, flag, or classification | Open-ended text, summaries, or images |
| Typical failure | A missed or false detection | Fabrication, omission, or inconsistent answers |
| Premarket evaluation | Performance on a fixed test set | Benchmarks across tasks plus clinical confirmation |
| Change over time | Locked, or managed through a change control plan | Underlying foundation models may be updated by third parties |
| User interaction | An alert or number to interpret | A conversation that shapes clinical reasoning |
| Postmarket signal | Measurable drift in accuracy | Content errors that are harder to detect at scale |
What the FDA Published
The FDA's public list of authorised AI-enabled devices now runs to well over a thousand products, and reports put the figure above 1,600. The overwhelming majority are in imaging and signal analysis, producing outputs a clinician can check against a defined standard. Generative products, such as tools that draft clinical notes, summarise records, answer clinical questions, or generate patient-facing explanations, have largely sat outside that list or been positioned to avoid device classification. The agency's digital health advisory committee examined generative AI devices in 2024, and the new discussion paper builds on that work.
The paper does four things. It proposes assessing risk along two axes rather than one, recognising that a generative tool's risk depends on more than its intended clinical area. It outlines competency assessment as the premarket evidence model. It explores how postmarket monitoring should work when errors are content rather than numbers. And it asks targeted questions about foundation models and agentic AI. Manufacturers should read the definitions in the paper itself, because the framing of those axes will shape how products are positioned.
Why Generative AI Medical Devices Are Different
A classic diagnostic algorithm can be wrong, but it is wrong in a bounded way. It misses a nodule or flags one that is not there, and both can be counted against a reference standard. Generative systems are wrong in open-ended ways, and the same input can produce different outputs on different runs.
How Generative AI Medical Devices Fail
- Fabrication. A summary includes a medication, finding, or history item that is not in the source record.
- Omission. A critical allergy, contraindication, or abnormal result is left out of an otherwise accurate summary.
- Inconsistency. The same question asked twice produces materially different answers.
- Prompt sensitivity. Small changes in wording lead to large changes in output.
- Upstream model change. An update to the foundation model alters behaviour without any change by the manufacturer.
- Automation bias. Fluent, confident text encourages clinicians to accept it with less scrutiny than they would a junior colleague's work.
See How Consistently Models Answer
Run the same clinical research question across six AI models and see where their answers diverge.
Try Talkory FreeCompetency Assessment in Plain Terms
The idea behind competency assessment is closer to how we judge a clinician than how we validate a lab test. Rather than one accuracy figure on one dataset, a manufacturer would show that the system performs acceptably across the range of tasks it is meant to do, using structured non-clinical benchmarks, and then confirm in a clinical setting that the benchmarked performance holds up with real users and real data.
In practice that suggests a different evaluation programme from the one most device teams are used to. It means defining the tasks the product performs, building test sets that cover typical and difficult cases, writing scoring rubrics that capture fabrication and omission as well as overall quality, using qualified clinical reviewers, measuring variation across repeated runs, and then running a clinical study that checks the benchmarks predict real-world performance. None of this is trivial, and much of it is expensive to do retrospectively.
Preparing Now: Seven Steps
- Write a precise intended use. Which tasks, which users, which settings, and what the output is used for.
- Map your risk on both axes. Use the paper's framing to understand where your product is likely to sit.
- Build a task-based benchmark. Cover routine and edge cases, with rubrics for fabrication, omission, and safety-relevant errors.
- Measure run-to-run variation. Report how much outputs vary for the same input, not just average quality.
- Plan for foundation model change. Decide how you will detect, assess, and control updates you do not make yourself.
- Design postmarket monitoring for content errors. Sampling, clinician feedback channels, and defined error categories.
- Submit comments or engage early. The comment window and pre-submission meetings are opportunities to shape expectations.
Pros and Cons of the Proposed Approach
- Pro: a clearer path to market. A defined evidence model reduces uncertainty for developers who want to do this properly.
- Pro: evaluation that fits the technology. Task-based competency testing captures failures a single accuracy number would miss.
- Pro: attention to foundation models. Third-party model dependency is a real risk that deserves explicit treatment.
- Con: higher evidence costs. Benchmarks, expert reviewers, and clinical confirmation add time and expense.
- Con: moving definitions. Until guidance is final, manufacturers must plan against a framework that may change.
- Con: monitoring is hard. Detecting content errors after deployment is far harder than tracking a performance metric.
Real Scenarios Worth Thinking Through
These scenarios are illustrative, showing how generative AI medical devices play out in practice rather than presented as verified case studies.
A company building a discharge summary tool reports high overall quality scores from clinician reviewers. When it adds a rubric for omissions, it finds that a small share of summaries leave out a medication change. The average score hid the safety-relevant failure, which is exactly what task-based competency testing is designed to expose.
A manufacturer relies on a third-party foundation model. The provider releases an update that improves general fluency but subtly changes how lab values are reported. Because the manufacturer runs its benchmark on every model version, the change is caught before it reaches clinicians.
A start-up positions its clinical question-answering tool as non-device decision support. After reading the discussion paper, its regulatory lead concludes that the way clinicians actually use it may not support that position, and the company begins building device-grade evidence early rather than late.
Evaluating Models on Sensitive Data?
Enterprise plans cover private deployment, custom data residency, dedicated infrastructure, and an SLA.
Talk to Enterprise Sales“After testing multiple AI models on coding, research, and business prompts, combined outputs produced more reliable results than any single model.” Internal multi-model evaluation, Talkory research team.
Foundation Models, Agents, and the Wider Picture
Two questions in the paper deserve particular attention. The first is foundation models. Most generative devices will be built on models the manufacturer did not train and cannot fully inspect, which shifts the focus to supplier agreements, version control, and evaluation of every update. The second is agentic AI: systems that take actions, not just produce text. An agent that orders a test or updates a record raises risk in a way a drafting tool does not.
Manufacturers selling in Europe face a parallel track, since AI in regulated medical devices falls within the EU's high-risk category, with timing still in flux as we covered in EU AI Act compliance after the omnibus delay. For AI used in development rather than in the device itself, the expectations look different again, covered in AI in clinical trials and validating GenAI in pharma.
Why Talkory Wins
Talkory is not a medical device and is not intended for clinical decisions. Where it helps is the research and evaluation work that sits behind a regulatory strategy. It runs the same question across GPT, Claude, Gemini, Grok, Perplexity Sonar, and Kimi K3 in one pass, which makes inconsistency between models immediately visible. Regulatory and evaluation teams can use that to test how well different models handle a task before choosing a foundation model, to find the questions where model answers diverge most and therefore belong in a benchmark, and to cross-check their reading of a regulatory question before taking it to counsel or the agency.
Final Verdict
Generative AI medical devices are moving from a regulatory grey zone toward a defined evidence model. The FDA's discussion paper points to two-axis risk assessment, competency assessment that pairs benchmarks with clinical confirmation, monitoring designed for content errors, and explicit controls for foundation models and agents. The details will change before anything is final. The direction is clear enough to act on: write a precise intended use, build task-based benchmarks with rubrics for fabrication and omission, measure variation, plan for upstream model change, and engage with the agency while the framework is still being shaped.
Find Where Models Disagree
Compare six AI models on the same question and see which answers belong in your benchmark.
Try Talkory FreeFrequently Asked Questions
What did the FDA publish on generative AI medical devices?
In August the FDA released a discussion paper, Considerations for the Regulation of Generative AI-Enabled Medical Devices, requesting public feedback. It covers risk assessment, premarket competency assessment, postmarket monitoring, foundation models, and agentic AI. Comments are due by October 19.
Is the FDA discussion paper binding guidance?
No. It is a discussion paper with a request for feedback, not final or draft guidance. It signals the agency's current thinking and the kind of evidence it is likely to expect, so manufacturers should plan with it in mind.
What is competency assessment for AI devices?
It is the proposed premarket evidence model combining non-clinical benchmarking across the tasks a system performs with clinical confirmation that performance holds in real use. It resembles assessing a clinician's competence more than validating a single lab test.
How do generative AI medical devices fail differently?
They can fabricate information, omit critical details, give inconsistent answers to the same input, react strongly to wording changes, and change behaviour when an underlying foundation model is updated. These failures are harder to measure than classification errors.
How should manufacturers prepare now?
Define a precise intended use, map risk using the paper's framing, build task-based benchmarks with rubrics for fabrication and omission, measure run-to-run variation, plan controls for foundation model updates, and design postmarket monitoring for content errors.
Get 5 AI perspectives on this topic
Talkory runs your question through GPT, Claude, Gemini, Grok, Sonar & Kimi K3 simultaneously, then cross-checks the answers.