The US Food and Drug Administration (FDA) issued a discussion paper on generative artificial intelligence-enabled medical devices on August 18, 2026, seeking feedback on risk assessment, evaluation before marketing, monitoring after release, foundation models and agentic systems. The consultation gives hospitals and developers a new set of questions about systems that can do more than generate an answer.

The FDA paper is for discussion only, is neither draft nor final guidance, and does not announce a policy change or a future evidence requirement. Its proposals apply to the US regulator's work on generative AI-enabled medical devices, not to every health care agent or every country.

A clinical team reviews a medical AI agent's task sequence and permission checkpoints at a workstation (illustrative image)
A medical AI agent links a task to data, external tools and a human decision point. (Illustrative image)

An agent can change a workflow, not merely draft an answer

The FDA defines agentic AI systems as generative AI-enabled systems that autonomously plan and execute multi-step tasks, use external tools or take actions across a sequence. The paper lists care coordination, clinical documentation, patient outreach and workflow support as current health care contexts, while noting that some of those functions may fall outside US device oversight.

A chatbot can stop after producing text. An agent may retrieve a record, select a tool, prepare a message, write to another system or trigger the next task when its permissions allow it. An error can therefore move from an inaccurate sentence into a record, notification or clinical workflow before another person sees it.

The action matters as much as the model. A scheduling draft and an autonomous change to a medical device can use similar planning technology but create different consequences when they fail. Intended use, user, data access, tool access and the reversibility of each action define the operational boundary.

Task position Example capability Control record
Prepare Read permitted records and produce a summary or draft Data sources, read permissions, provenance and editable output
Coordinate Propose a referral packet, appointment message or next task Allowed tools, recipient rules, human approval and timeout route
Act Write to a clinical record, start an order or control another device Case-specific authorization, audit trail, stop authority and recovery path

This is an APPI News editorial grouping, not an FDA classification or a claim that every example is a medical device. Product status depends on the intended function and the rules of each country. A broader APPI News deployment framework separates intended use, evidence, interoperability, human control and monitoring; agentic systems add a sequence of tool calls and actions that must stay inside those controls.

First safeguard: bind every tool to a permission

An agent's permission record needs more detail than a label such as “clinical assistant.” It can list each data source and tool, the operations allowed, prohibited fields, affected users, time limits and conditions that require approval. Read, propose, write and approve are separate permissions even when one interface presents them as a single task.

The FDA's possible agentic competency asks whether a system stays within its intended use, uses tools accurately, recognizes erroneous tool outputs, observes human checkpoints before irreversible or high-consequence actions, handles tool failures and resists prompt injection. The appendix presents a possible evaluation element for public discussion, not an adopted US test standard.

A permission test can start with a normal task and then remove a record, return a conflicting result, deny a tool call or place an instruction inside retrieved content. The record should show whether the agent stops, requests missing information, transfers the task or attempts an action outside its scope. A successful demonstration under ideal inputs does not answer those failure cases.

A digital health workflow separates data input, tool access, permissions and audit events (illustrative image)
Data exchange, tool permission and authorization are separate layers in an agent workflow. (Illustrative image)

HL7 International's Fast Healthcare Interoperability Resources (FHIR) Release 5 defines modular resources and interfaces for electronic health-information exchange. It can provide a common structure for records moving between a hospital system and an agent, but a shared structure does not decide what the agent may read or change.

HL7 says a production FHIR system needs a separate security subsystem for authentication, access-control decisions and audit logging. The implementation must also validate that the patient, code, unit, timestamp and record context mean the same thing on both sides of the interface.

Second safeguard: name the person before a consequential action

World Health Organization guidance says humans should remain in control of health systems and medical decisions and that privacy and confidentiality require protection. That principle becomes a testable control only when a deployment record names the person with authority at each decision point.

The International Medical Device Regulators Forum's 2025 principles call for assessing the human-AI team in the intended environment, including user understanding, overreliance, the level of autonomy and foreseeable misuse. A confirmation button alone does not establish that review works under clinical conditions.

The handoff record can specify what evidence the reviewer sees, whether the reviewer may edit or reject the output, how disagreement is recorded and what happens when nobody responds. It also needs an escalation route for missing, contradictory or unsupported input. The person at the checkpoint must have enough time, information and authority to change the result.

Joint 2024 principles from the US FDA, Health Canada and the United Kingdom's Medicines and Healthcare products Regulatory Agency say medical-device transparency should cover intended use, workflow effects, performance, risks, limitations, data gaps and maintenance across the lifecycle. They also identify update notifications and targeted information during high-risk workflow steps as good practices, not a single law that applies worldwide.

The same record should separate the hospital's release decision from a supplier's product claim and a clinician's decision on an individual case. APPI News has mapped those positions into records for intended use, data, human workflow, traceability, change control and incidents. An agent adds another question at every step: which actor authorized the tool call that followed?

Hospital technology and clinical staff review AI governance, validation and human approval controls (illustrative image)
A human checkpoint needs a named reviewer, the evidence needed for review and a route for rejecting the action. (Illustrative image)

Third safeguard: monitor the task chain and preserve recovery

Monitoring an agent requires more than sampling its final prose. The record can connect the input, planning step, retrieved sources, tool calls, model and software versions, approval event, action taken and later correction. That sequence helps an investigator distinguish a model error from a bad data mapping, failed tool or unauthorized action.

The FDA paper identifies periodic benchmarking, sample-based review by independent clinicians and performance-degradation monitoring as possible approaches after release. It also asks how manufacturers can detect and respond when a third-party foundation-model developer changes refusal behavior, content policies, output format, version control or another safety-related behavior.

A monitoring plan needs a baseline, review schedule, alert conditions and an investigation owner. It also needs a named authority who can narrow permissions, suspend the agent or restore an earlier configuration. Recovery may mean reverting a model, disabling one tool or returning the whole task to a manual workflow, so each option needs its own test.

Every release should preserve the version and acceptance evidence that preceded it. APPI News's medical AI validation guide connects representative testing, workflow checks, monitoring signals and predefined pause and rollback routes. For an agent, that evidence must also cover the tools and orchestration rules used by the released version.

A hospital monitoring screen displays an AI permission matrix, review alerts and a rollback control (illustrative image)
Monitoring must lead to a decision to continue, restrict, suspend or restore the service. (Illustrative image)

One acceptance file connects the three safeguards

A predeployment file can connect the intended task, permission matrix, representative test cases, human checkpoints, audit design, monitoring plan and recovery procedure. Each record must identify the same users, data sources, tools, model version and allowed actions. A mismatch means the evidence describes a different system from the one being deployed.

  • Scope record: intended users, setting, input, output, allowed actions, prohibited actions and fallback.
  • Permission record: data sources, tools, read and write operations, approval conditions and access revocation.
  • Evaluation record: normal cases, missing and conflicting inputs, bad tool results, attempted scope expansion, prompt injection and outages.
  • Human-workflow record: reviewer, evidence shown, response time, accept or reject authority, escalation and unresolved cases.
  • Operations record: versions, logs, monitoring measures, alert conditions, investigation owner, suspension and recovery.

This file is an APPI News synthesis of the cited documents, not a certification program, regulatory approval or legal test. Authorization, professional duties, privacy requirements and incident reporting must be checked in every country where a system operates.

Published data do not yet isolate agent deployment

The FDA discussion paper describes agentic systems and asks how their additional risks should affect evaluation. It does not report how many US hospitals use them, establish that any named product is safe, or show that one control design works across health systems.

APPI News could not find a comparable published cross-country dataset showing how many hospitals permit AI agents to execute clinical actions. Available international surveys often group diagnostic systems, chatbots, administrative tools and other AI applications together, so they cannot establish adoption of agent-specific permission, approval and recovery controls.

The three safeguards preserve different evidence: permissions show what the agent was allowed to do, human checkpoints show who authorized a consequential action, and monitoring shows what happened after release. A hospital can trace an agentic workflow only when those records describe the same task chain and the same deployed version.