The US Food and Drug Administration (FDA) issued a discussion paper on generative artificial intelligence-enabled medical devices on August 18, 2026, seeking feedback on risk assessment, evaluation before marketing, monitoring after release, foundation models and agentic systems. The consultation gives hospitals and developers a new set of questions about systems that can do more than generate an answer.
The FDA paper is for discussion only, is neither draft nor final guidance, and does not announce a policy change or a future evidence requirement. Its proposals apply to the US regulator's work on generative AI-enabled medical devices, not to every health care agent or every country.
An agent can change a workflow, not merely draft an answer
The FDA defines agentic AI systems as generative AI-enabled systems that autonomously plan and execute multi-step tasks, use external tools or take actions across a sequence. The paper lists care coordination, clinical documentation, patient outreach and workflow support as current health care contexts, while noting that some of those functions may fall outside US device oversight.
A chatbot can stop after producing text. An agent may retrieve a record, select a tool, prepare a message, write to another system or trigger the next task when its permissions allow it. An error can therefore move from an inaccurate sentence into a record, notification or clinical workflow before another person sees it.
The action matters as much as the model. A scheduling draft and an autonomous change to a medical device can use similar planning technology but create different consequences when they fail. Intended use, user, data access, tool access and the reversibility of each action define the operational boundary.
| Task position | Example capability | Control record |
|---|---|---|
| Prepare | Read permitted records and produce a summary or draft | Data sources, read permissions, provenance and editable output |
| Coordinate | Propose a referral packet, appointment message or next task | Allowed tools, recipient rules, human approval and timeout route |
| Act | Write to a clinical record, start an order or control another device | Case-specific authorization, audit trail, stop authority and recovery path |
This is an APPI News editorial grouping, not an FDA classification or a claim that every example is a medical device. Product status depends on the intended function and the rules of each country. A broader APPI News deployment framework separates intended use, evidence, interoperability, human control and monitoring; agentic systems add a sequence of tool calls and actions that must stay inside those controls.
First safeguard: bind every tool to a permission
An agent's permission record needs more detail than a label such as “clinical assistant.” It can list each data source and tool, the operations allowed, prohibited fields, affected users, time limits and conditions that require approval. Read, propose, write and approve are separate permissions even when one interface presents them as a single task.
The FDA's possible agentic competency asks whether a system stays within its intended use, uses tools accurately, recognizes erroneous tool outputs, observes human checkpoints before irreversible or high-consequence actions, handles tool failures and resists prompt injection. The appendix presents a possible evaluation element for public discussion, not an adopted US test standard.
A permission test can start with a normal task and then remove a record, return a conflicting result, deny a tool call or place an instruction inside retrieved content. The record should show whether the agent stops, requests missing information, transfers the task or attempts an action outside its scope. A successful demonstration under ideal inputs does not answer those failure cases.
HL7 International's Fast Healthcare Interoperability Resources (FHIR) Release 5 defines modular resources and interfaces for electronic health-information exchange. It can provide a common structure for records moving between a hospital system and an agent, but a shared structure does not decide what the agent may read or change.
HL7 says a production FHIR system needs a separate security subsystem for authentication, access-control decisions and audit logging. The implementation must also validate that the patient, code, unit, timestamp and record context mean the same thing on both sides of the interface.
Second safeguard: name the person before a consequential action
World Health Organization guidance says humans should remain in control of health systems and medical decisions and that privacy and confidentiality require protection. That principle becomes a testable control only when a deployment record names the person with authority at each decision point.
The International Medical Device Regulators Forum's 2025 principles call for assessing the human-AI team in the intended environment, including user understanding, overreliance, the level of autonomy and foreseeable misuse. A confirmation button alone does not establish that review works under clinical conditions.
The handoff record can specify what evidence the reviewer sees, whether the reviewer may edit or reject the output, how disagreement is recorded and what happens when nobody responds. It also needs an escalation route for missing, contradictory or unsupported input. The person at the checkpoint must have enough time, information and authority to change the result.
Joint 2024 principles from the US FDA, Health Canada and the United Kingdom's Medicines and Healthcare products Regulatory Agency say medical-device transparency should cover intended use, workflow effects, performance, risks, limitations, data gaps and maintenance across the lifecycle. They also identify update notifications and targeted information during high-risk workflow steps as good practices, not a single law that applies worldwide.
The same record should separate the hospital's release decision from a supplier's product claim and a clinician's decision on an individual case. APPI News has mapped those positions into records for intended use, data, human workflow, traceability, change control and incidents. An agent adds another question at every step: which actor authorized the tool call that followed?
Third safeguard: monitor the task chain and preserve recovery
Monitoring an agent requires more than sampling its final prose. The record can connect the input, planning step, retrieved sources, tool calls, model and software versions, approval event, action taken and later correction. That sequence helps an investigator distinguish a model error from a bad data mapping, failed tool or unauthorized action.
The FDA paper identifies periodic benchmarking, sample-based review by independent clinicians and performance-degradation monitoring as possible approaches after release. It also asks how manufacturers can detect and respond when a third-party foundation-model developer changes refusal behavior, content policies, output format, version control or another safety-related behavior.
A monitoring plan needs a baseline, review schedule, alert conditions and an investigation owner. It also needs a named authority who can narrow permissions, suspend the agent or restore an earlier configuration. Recovery may mean reverting a model, disabling one tool or returning the whole task to a manual workflow, so each option needs its own test.
Every release should preserve the version and acceptance evidence that preceded it. APPI News's medical AI validation guide connects representative testing, workflow checks, monitoring signals and predefined pause and rollback routes. For an agent, that evidence must also cover the tools and orchestration rules used by the released version.
One acceptance file connects the three safeguards
A predeployment file can connect the intended task, permission matrix, representative test cases, human checkpoints, audit design, monitoring plan and recovery procedure. Each record must identify the same users, data sources, tools, model version and allowed actions. A mismatch means the evidence describes a different system from the one being deployed.
- Scope record: intended users, setting, input, output, allowed actions, prohibited actions and fallback.
- Permission record: data sources, tools, read and write operations, approval conditions and access revocation.
- Evaluation record: normal cases, missing and conflicting inputs, bad tool results, attempted scope expansion, prompt injection and outages.
- Human-workflow record: reviewer, evidence shown, response time, accept or reject authority, escalation and unresolved cases.
- Operations record: versions, logs, monitoring measures, alert conditions, investigation owner, suspension and recovery.
This file is an APPI News synthesis of the cited documents, not a certification program, regulatory approval or legal test. Authorization, professional duties, privacy requirements and incident reporting must be checked in every country where a system operates.
Published data do not yet isolate agent deployment
The FDA discussion paper describes agentic systems and asks how their additional risks should affect evaluation. It does not report how many US hospitals use them, establish that any named product is safe, or show that one control design works across health systems.
APPI News could not find a comparable published cross-country dataset showing how many hospitals permit AI agents to execute clinical actions. Available international surveys often group diagnostic systems, chatbots, administrative tools and other AI applications together, so they cannot establish adoption of agent-specific permission, approval and recovery controls.
The three safeguards preserve different evidence: permissions show what the agent was allowed to do, human checkpoints show who authorized a consequential action, and monitoring shows what happened after release. A hospital can trace an agentic workflow only when those records describe the same task chain and the same deployed version.
Sources and further reading
- FDA Seeks Public Feedback to Inform Regulatory Approach for Generative AI-Enabled Medical Devices(US Food and Drug Administration)August 18, 2026
- Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback(US Food and Drug Administration)Discussion paper, August 2026
- Good machine learning practice for medical device development: Guiding principles(International Medical Device Regulators Forum)Final document, January 27, 2025
- Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles(US Food and Drug Administration, Health Canada and the United Kingdom's Medicines and Healthcare products Regulatory Agency)June 2024
- WHO issues first global report on artificial intelligence in health and six guiding principles(World Health Organization)June 28, 2021
- FHIR Overview(HL7 International)
- FHIR Security(HL7 International)