FDA, Before Competency Comes Necessity: The Hidden Liability of Integrating Generative AI into Medical Devices
The Missing Threshold in the FDA's Generative AI Medical Device Framework
BBIU Institutional Analysis | Regulatory Architecture, Data Governance, and Systemic Risk | August 2026
The U.S. Food and Drug Administration is asking how generative artificial intelligence-enabled medical devices should be classified, evaluated, and monitored. BBIU's analysis identifies a logically prior question:
Why does the intended medical function require generative AI, and why must that AI be integrated into the device?
That distinction matters. A device can be technically complex, highly automated, and clinically consequential without requiring artificial intelligence. It can also use narrow predictive machine learning without requiring a generative foundation model. If bounded software, deterministic control logic, or a separable interpretive service can perform the intended function, integrating GenAI introduces additional variability, infrastructure dependency, cybersecurity exposure, and patient-data transmission that must be justified before competence is evaluated.
The regulatory sequence should therefore begin with necessity, not performance.
Critical Takeaways
The FDA's August 2026 publication is a discussion paper, not draft or final guidance, and it does not establish new evidentiary requirements. Public comments are due under docket FDA-2026-N-7874 by October 19, 2026.
The agency is considering a two-axis risk framework, competency-based premarket evaluation, clinical confirmation, postmarket monitoring, Foundation Model Device Master Files, and additional regulatory considerations for agentic systems.
The paper recognizes many obvious limitations, including benchmark contamination, synthetic-data bias, foundation-model opacity, model changes initiated by third parties, and the risk of diffusing manufacturer accountability.
What remains insufficiently addressed is the threshold question: whether GenAI is necessary and whether it must be embedded in the regulated product.
An integrated architecture can leave the manufacturer responsible for the device while dependent on a safety-relevant model and infrastructure it does not fully control.
FDA authorization would not, by itself, establish compliance with HIPAA, FTC health-data rules, state privacy law, or the full cloud and data-retention architecture.
Once identifiable medical information is disclosed to an external AI infrastructure, deletion may reduce retained copies but cannot restore the previous state of confidentiality.
For many interpretive uses, a safer default may be a standardized, patient-controlled source-data package with AI maintained as a separable, user-selectable layer.
1. What the FDA Is Actually Considering
On August 18, 2026, the FDA's Center for Devices and Radiological Health released Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback. The document is explicitly exploratory: it is not guidance, does not communicate final or proposed evidentiary expectations, and does not determine whether every approach discussed falls within the FDA's existing legal authority.
Nevertheless, the paper reveals the architecture CDRH is examining.
First, the FDA proposes a possible two-axis heuristic. One axis measures the degree and independence of the device's activity, moving from non-directive information toward supervised and then fully autonomous action. The other measures the severity of harm that could result from relying on an incorrect output. The FDA also recognizes that risk may evolve across multi-turn conversations and may differ depending on whether the user is a patient, generalist physician, or specialist.
Second, CDRH is considering a competency-based premarket framework inspired at a high level by physician evaluation. The final user-facing product—not only the underlying foundation model—could undergo:
Non-clinical device benchmarking, assessing safety behavior, clinical proficiency, communication, robustness, reproducibility, subgroup performance, and agentic capabilities where applicable.
Clinical confirmation, potentially using retrospective cases, shadow deployment, standardized patients, independent clinician adjudication, prospective studies, or randomized trials depending on risk.
Third, the FDA is considering whether stronger postmarket monitoring could sometimes justify accepting greater premarket uncertainty. Possible tools include periodic re-benchmarking, clinician review of sampled real-world interactions, and performance-degradation monitoring. The agency also connects this approach to its existing framework for Predetermined Change Control Plans, which allows sponsors to describe planned modifications and the methods used to develop, validate, and implement them.
Finally, the paper addresses two structural dependencies. For devices built on third-party models, it explores voluntary Foundation Model Device Master Files containing model cards, system cards, limitations, safety controls, subgroup performance, update commitments, and audit-log information. For agentic systems, it asks how autonomous multi-step planning, tool use, and reduced opportunities for human review should affect regulatory expectations.
These are substantive questions. But they begin after the architectural choice has already been made.
2. The Missing Threshold: Why GenAI?
Medical-device regulation should not treat technical complexity as evidence that generative AI is necessary.
Many critical medical functions can be performed through deterministic software, mathematical models, signal-processing pipelines, predefined clinical rules, or conventional feedback control. Such systems can be sophisticated without generating open-ended outputs. Narrow machine learning may also be useful for classification, segmentation, anomaly detection, or prediction without requiring a general-purpose or multimodal foundation model.
Generative behavior is different. It becomes potentially valuable when the intended function genuinely requires open-ended inputs, variable outputs, conversational interaction, synthesis across heterogeneous information, or flexible multi-step reasoning. Those same properties, however, make behavior harder to bound, reproduce, and exhaustively test.
The manufacturer should therefore have to establish more than adequate GenAI performance. It should have to show that the clinical benefit cannot reasonably be obtained through a simpler and more controllable architecture.
BBIU proposes evaluating this choice through a Least-Generative Principle, formalized later in this analysis. Under that principle, the sponsor would need to:
Define the unmet clinical need.
Identify the simplest bounded architecture capable of addressing it.
Compare the GenAI architecture against that alternative.
Demonstrate incremental clinical benefit.
Quantify the additional technological, privacy, and cybersecurity risk.
Explain why the generative function must be integrated into the device rather than offered as a separable service.
The FDA repeatedly invokes least-burdensome regulation. The missing counterpart is a preference for the least unpredictable technology capable of accomplishing the intended medical purpose.
3. Integration Is a Separate Decision From Using AI
Even when a generative model adds value, it does not follow that the model must become part of the medical device.
Integration may be justified when separation would materially compromise:
real-time access to device-native signals;
validated interaction with hardware;
latency-sensitive clinical response;
local or offline operation;
safety interlocks;
end-to-end clinical effectiveness; or
the ability to verify the complete system under its intended conditions of use.
For many explanatory and interpretive functions, however, those conditions may not apply. Summarizing results, translating technical information, preparing questions for a physician, or explaining trends could potentially remain outside the core device. A patient or clinician may already have access to a preferred model that is more current, multilingual, or integrated with other authorized information sources.
The relevant question is therefore not only why GenAI, but why integrated GenAI.
That distinction creates two different regulatory architectures:
Integrated architecture: medical device → manufacturer-selected model and cloud stack → generated interpretation or action
Separated architecture: medical device → standardized patient-controlled data package → user-selected AI → optional clinical review
The first concentrates product validation but also concentrates exposure and dependency. The second increases user choice and model portability but requires strict separation between personal interpretation and regulated clinical action.
4. Manufacturer Accountability Without Complete Control
When a sponsor incorporates a third-party foundation model into a medical device, it remains responsible for demonstrating the safety and effectiveness of its own product. The FDA discussion paper states that a Foundation Model Master File would not authorize the underlying model for a medical intended use and would not replace the sponsor's device-specific evidence.
Yet the sponsor may not fully control the component on which its product depends.
The manufacturer can control its user interface, system prompts, retrieval sources, acceptance filters, deterministic post-processing, and some model-version policies. It may negotiate update notifications, restrict retention, conduct re-benchmarking, and introduce human review. Those measures can reduce risk. They do not give the sponsor complete visibility into or control over:
original training-data provenance;
internal model architecture and evaluation;
emergent failure modes;
provider-level refusal behavior and guardrails;
output variability across semantically similar inputs;
provider-side logging, retention, and subprocessors;
upstream cybersecurity vulnerabilities;
API outages, withdrawal, or replacement; or
future changes initiated by the model developer.
This is not an argument that no medical technology can depend on third parties. Medical devices routinely incorporate external software, hardware, and cloud services. The distinctive issue is the combination of open-ended behavior, limited transparency, continuous upstream change, and clinical reliance.
The resulting asymmetry is central:
The manufacturer retains primary regulatory responsibility for the integrated medical product while lacking complete technical control over a safety-relevant component operated by another company.
Contracts can allocate obligations and liability between the device manufacturer, model provider, cloud provider, healthcare institution, and other participants. They cannot eliminate the sponsor's obligation to support the safety and effectiveness of the regulated product.
5. FDA Authorization Would Not Approve the Entire Data Architecture
An FDA marketing authorization addresses the regulated device. It does not automatically establish compliance across every legal regime implicated by the product's data flows.
An integrated GenAI device may involve:
medical device → manufacturer → cloud platform → foundation-model provider → monitoring infrastructure → subcontractors
When a HIPAA-covered entity or business associate uses a cloud provider to create, receive, maintain, or transmit electronic protected health information on its behalf, HHS explains that the cloud provider generally becomes a business associate—even when it cannot view encrypted information. A compliant Business Associate Agreement is required, and using such a provider without one may violate HIPAA. HHS also requires downstream subcontractors handling PHI on behalf of a business associate to accept corresponding restrictions and safeguards. See the official HHS guidance on HIPAA and cloud computing and business associates.
Outside HIPAA, the absence of HIPAA coverage does not mean the absence of regulation. The FTC's amended Health Breach Notification Rule expressly reaches many health apps, connected devices, and related products that are not covered by HIPAA. State privacy and breach-notification statutes may also apply, while international deployment may trigger additional data-protection and cross-border-transfer requirements.
The FDA's GenAI discussion paper does not substantively address HIPAA, patient-data privacy, or the broader cross-agency data-governance stack. That omission does not mean the FDA ignores cybersecurity generally. The agency maintains separate medical-device cybersecurity guidance, including recommendations for cybersecurity design, labeling, and premarket documentation. The problem is fragmentation: a GenAI device can receive FDA authorization while still depending on unresolved privacy, cloud, contractual, and cross-border compliance decisions.
6. Medical-Data Disclosure Is Not Fully Reversible
Once identifiable medical information is disclosed to an external AI infrastructure, the previous state of confidentiality cannot be recreated.
Deletion remains important, but it is not equivalent to reversal. HHS requires Business Associate Agreements to provide for the return or destruction of PHI at termination where feasible. Where destruction is not feasible, protections must continue and further uses must remain limited. That qualification is important: even under HIPAA-governed relationships, complete destruction may not always be operationally possible.
The data lifecycle may include:
prompts and generated outputs;
interaction histories;
system and security logs;
backups;
human-review records;
retrieval indexes and embeddings;
model-evaluation datasets;
downstream processors; and
clinical or behavioral inferences derived from the original information.
NIST's Generative AI Profile identifies privacy risks associated with model memorization and sensitive-data exposure. Its separate Adversarial Machine Learning taxonomy addresses reconstruction, membership inference, model leakage, poisoning, and other attacks across the AI lifecycle.
Even if every known stored copy were later removed, deletion cannot undo the fact that another entity accessed the information, processed it, generated inferences, or used those inferences in a decision.
BBIU therefore proposes a Data Irreversibility Principle:
Once identifiable health information leaves the controlled device boundary, its disclosure should be treated as irreversible for purposes of architectural risk assessment, regardless of later deletion commitments.
This principle does not prohibit cloud processing. It requires the necessity of the disclosure to be established before it occurs.
7. Cybersecurity Expands Beyond the Device Boundary
Integrating a generative model adds attack surfaces that are not limited to conventional device software.
NIST identifies prompt injection and data poisoning among the distinctive security risks of generative AI. In a medical-device architecture, additional pathways may include compromised API credentials, manipulated retrieval sources, malicious content embedded in connected records, unauthorized access to prompts and outputs, provider-side supply-chain compromise, and adversarial instructions directed at agentic systems.
The FDA's current cybersecurity framework appropriately treats interoperability as part of end-to-end device security. But GenAI extends the relevant boundary beyond the manufacturer's application into the model, orchestration layer, retrieval environment, cloud infrastructure, monitoring system, and subprocessor chain.
No manufacturer can guarantee absolute security, and no medical device is risk-free. The decisive question is different:
Does the incremental clinical benefit of the integrated GenAI function justify the additional and less controllable attack surface?
If the same clinical function can be delivered locally, deterministically, or through a patient-controlled external service, the additional exposure becomes harder to defend.
8. From Device Risk to Common-Mode Infrastructure Risk
The FDA's two-axis framework focuses on the activity performed by a device function and the consequence of relying on an incorrect output. BBIU considers that framework incomplete for foundation-model-dependent products.
A third dimension is necessary: scale and concentration of dependency.
Multiple authorized devices may depend on the same model family, API, cloud provider, retrieval component, or safety layer. A single upstream update, outage, vulnerability, or behavioral change could therefore affect several products and healthcare institutions simultaneously. The individual device may remain unchanged from the manufacturer's perspective while its real-world behavior changes because a shared upstream component changed.
This is common-mode infrastructure risk. It is not captured by measuring the severity of one incorrect output in one device.
The FDA paper recognizes the underlying mechanism when it asks how manufacturers should detect and respond to third-party model changes. BBIU's inference is that the regulatory unit may sometimes need to extend beyond the individual device toward dependency concentration, propagation speed, installed exposure, and coordinated suspension or rollback capacity.
9. Competency Without a True Residency Stage
The FDA's physician-credentialing analogy is useful but incomplete.
Physicians do not move from examination directly to independent practice. Their pathway includes supervised clinical exposure, progressive responsibility, institutional privileging, professional accountability, and continuing evaluation.
The discussion paper includes clinical confirmation methods such as retrospective evaluation, shadow deployment, clinician adjudication, and prospective trials. Shadow deployment is valuable, but it does not reproduce supervised clinical responsibility because the system's outputs are hidden and do not affect care. Postmarket monitoring is also not equivalent to supervision: it may identify failures after exposure has occurred.
For higher-risk integrated or agentic systems, a more faithful competency model would include staged authorization:
limited initial deployment;
capped patient exposure;
mandatory human review;
restricted autonomy;
predefined escalation and suspension thresholds;
outcome-linked monitoring; and
progressive expansion only after demonstrated performance.
If greater premarket uncertainty is accepted, graduated deployment becomes more—not less—important.
10. A Patient-Controlled Alternative
For many interpretive functions, the safer default may be to separate the regulated data-producing device from the AI interpretation layer.
Instead of automatically sending patient information to a manufacturer-selected model, the device could provide a standardized, portable, and verifiable source-data package controlled by the patient or authorized user.
The architecture would be:
medical device → standardized patient-controlled source-data package → user-selected AI → optional clinical review
This alternative does not treat undocumented sensor output as sufficient. "Raw" should mean source-level clinical evidence without a newly imposed AI interpretation—not data stripped of the metadata necessary to understand it.
Depending on the device and intended use, the package could include:
original diagnostic images;
digital pathology whole-slide images;
ECG, EEG, or other physiological waveforms;
laboratory measurements;
vital-sign time series;
structured medications, allergies, and relevant history;
device-native sensor data;
acquisition and calibration parameters;
units, timestamps, reference ranges, and quality indicators;
device provenance, software version, and algorithm version; and
integrity verification through a digital signature or checksum.
The package should explicitly separate:
source measurements and observations;
existing professional reports;
manufacturer-generated calculations; and
AI-generated interpretations.
This separation allows the recipient to determine what was measured, what was calculated, what was concluded by a professional, and what was generated by an AI system.
11. Interoperability Makes the Alternative Technically Plausible
The required infrastructure would not need to be invented from zero.
DICOM is the international standard for medical images and related information. HL7 FHIR provides a standard for electronic healthcare data exchange, including resources for devices, observations, diagnostic reports, imaging studies, medications, and clinical records. The Office of the National Coordinator's United States Core Data for Interoperability defines standardized health-data classes and elements for nationwide exchange. ONC also requires standardized FHIR APIs using USCDI content for certified patient and population services, demonstrating that patient-facing, standards-based data exchange is already part of the U.S. health IT architecture.
The FDA likewise has longstanding interoperable medical-device guidance covering devices designed to exchange and use information with other medical and non-medical products.
The innovation would not be interoperability alone. It would be turning interoperability into an architectural constraint against unnecessary model lock-in: the patient receives usable, authenticated clinical data before deciding whether an external AI should process it.
12. Personal Interpretation Is Not the Same as Regulated Clinical Action
A patient-controlled architecture does not mean that any general-purpose LLM should be allowed to diagnose, prescribe, modify treatment, or control a device.
The external interface should be read-only by default. A patient may use a model for education, explanation, translation, or preparation of questions. If the AI performs a medical function—such as generating a diagnosis, treatment recommendation, dosage change, or device-control command—that AI function may itself fall within medical-device regulation depending on its intended use and statutory status.
If an external model can write instructions back to the device, the separation has ended. The model, interface, and hardware may need to be evaluated as a combined system.
Under a genuinely separated architecture, responsibilities become clearer:
the device manufacturer is responsible for the accuracy, integrity, and documentation of device-generated data;
the AI provider is responsible for representations and functions it offers;
the patient controls whether and where the data are transferred; and
the clinician determines whether an external interpretation should influence professional care.
This does not eliminate overlapping legal responsibility. It aligns primary responsibility more closely with the layer each party controls.
13. HIPAA Exposure Is Reduced, Not Automatically Preserved
A patient-controlled architecture can reduce manufacturer- or provider-initiated HIPAA exposure because the covered entity provides information to the individual rather than automatically transmitting it through a manufacturer-selected AI chain.
HHS explains that an app's facilitation of access to an individual's ePHI at the individual's request does not, by itself, create a business-associate relationship. The conclusion depends on the relationship between the app, covered entity, and other participants. See the HHS FAQ on patient-directed app access and Business Associate Agreements.
The protection gap appears after the patient transfers the data. HHS states in its guidance on health information stored on personal devices and apps that HIPAA generally does not protect information entered into or stored by a personal app when the app is not provided by a covered entity or its business associate. The data remain medically sensitive, but they may leave the HIPAA-regulated environment. The FTC Health Breach Notification Rule and other privacy laws may still apply.
The accurate formulation is therefore:
A patient-controlled architecture can reduce provider- or manufacturer-initiated HIPAA exposure, but it does not guarantee that HIPAA protection continues after the patient transfers the data to a non-covered AI provider.
The patient should see, before transfer:
the exact data being sent;
the recipient and model provider;
retention and deletion terms;
whether the data may be used for training or product improvement;
relevant subprocessors;
whether HIPAA applies; and
whether the model is intended or authorized for a clinical function.
Patient control does not make disclosure reversible. It prevents external AI exposure from becoming an unavoidable condition of using the core medical device.
14. The Commercial Incentive and Regulatory Contradiction
Why would a manufacturer integrate GenAI when a separable architecture is possible?
BBIU's inference is that integration may create substantial commercial value. It can allow the manufacturer to monetize interpretation, introduce recurring software revenue, differentiate otherwise comparable hardware, retain the direct relationship with the patient, collect additional usage data, and increase switching costs.
Those benefits should not be treated as evidence of improper conduct. They are rational product-strategy incentives. But they must be separated from clinical necessity.
The same integration also creates additional burdens:
continuing monitoring and revalidation;
dependence on an external model provider;
cybersecurity and privacy exposure;
contractual complexity;
possible model discontinuation or substitution;
incident and breach response;
common-mode risk; and
liability arising from outputs the manufacturer cannot completely control.
This produces the central commercial-regulatory contradiction:
The commercial value of integration comes from controlling the interpretation layer, while the principal regulatory risk comes from the manufacturer's inability to fully control that same layer.
The FDA should therefore distinguish the manufacturer's business rationale from the patient's clinical need.
15. Three Proposed Regulatory Principles
BBIU proposes three connected principles.
Least-Generative Principle
A medical device should use the least stochastic and least generative technology capable of safely and effectively performing its intended function.
Least-Exposure Principle
A medical device should not expose patient data to a generative AI infrastructure unless the manufacturer demonstrates that the exposure is necessary to achieve a clinically meaningful benefit that cannot be obtained through a less data-intensive architecture.
Open Device–AI Separation Principle
Medical devices should provide standardized, patient-controlled clinical data, while AI should remain a separable and user-selectable interpretation layer unless integration is clinically necessary for safety or effectiveness.
These principles do not prohibit integrated GenAI. They create a rebuttable presumption in favor of simpler technology, minimized disclosure, and model portability.
16. The Questions That Should Precede Competency Testing
Before determining whether an integrated GenAI-enabled device is competent, regulators and review committees should require the sponsor to answer:
What specific medical limitation requires generative rather than deterministic or narrow predictive technology?
What bounded alternatives were evaluated, and why were they insufficient?
What incremental clinical benefit does the GenAI function produce?
Why must the model be integrated into the device rather than remain separable?
What patient data must leave the controlled device boundary, and why?
Can the function operate locally or with a smaller data set?
Which entities create, receive, maintain, transmit, log, review, or derive information from the data?
Which model components and behaviors can the manufacturer control, observe, freeze, and roll back?
What happens if the foundation-model provider changes, withdraws, or compromises the model?
How concentrated is the dependency across other devices and healthcare institutions?
What staged-deployment and suspension mechanisms apply before unrestricted use?
Can the patient use the core device without accepting secondary AI processing?
Can the patient obtain a standardized, authenticated source-data package for use elsewhere?
Only after those questions are answered should competency benchmarking determine whether the chosen architecture performs adequately.
Conclusion: The Regulatory Object Is the Architecture
The FDA's discussion paper is a serious attempt to confront technologies that produce variable outputs, evolve over time, depend on opaque foundation models, and may operate with increasing autonomy. Its competency-based framework, clinical-confirmation options, postmarket monitoring concepts, and attention to third-party models provide a credible starting point.
But competence is not the first decision.
A GenAI-enabled medical device may perform well on benchmarks while still embodying an unnecessary architecture: a simpler system may accomplish the clinical function; the model may not need to be integrated; patient data may travel through avoidable infrastructure; and the manufacturer may remain accountable for behavior it cannot fully control.
The regulatory object should therefore be broader than the generated output. It should include the decision to use GenAI, the decision to integrate it, the data path required to operate it, the concentration of its external dependencies, and the patient's ability to refuse or separate the AI layer.
The FDA begins by asking how GenAI medical devices should be evaluated. The prior institutional question is more fundamental:
Should this medical function have been implemented through an integrated generative architecture in the first place?