The International Medical Device Regulators Forum (IMDRF) published final principles for planned medical-software changes on August 6, 2026. A predetermined change control plan identifies specific changes, an implementation and control protocol with prespecified criteria, and an assessment of their effects; covered changes remain within the software's original intended purpose.

The update framework completes a lifecycle that starts before a model receives data and continues after release. Medical AI, or the use of artificial intelligence in health care, is an umbrella for many tools rather than one product class, and a system's intended purpose and the rules in each market determine whether it enters a medical-device pathway.

A clinical team examines AI results and a patient-data workflow on a monitor (illustrative image)
A medical AI workflow begins with a defined purpose, input and user, not with a model name. (Illustrative image)

Medical AI covers several kinds of work

World Health Organization (WHO) guidance published in January 2024 groups health uses of large multimodal models into diagnosis and clinical care, patient-guided use, administration, education, and research and drug development. That list concerns generative models, while the wider medical AI field also includes image classifiers, risk models, signal analysis and systems that rank cases for review.

A model can produce a class, probability, ranking, measurement, summary or draft. Its place in the workflow decides who sees that output, what decision it may influence and what evidence is needed.

Function Typical inputs Output and user Control question
Clinical interpretation Images, waveforms or laboratory values A finding, measurement or score for trained staff Does the evidence cover the intended population, equipment and setting?
Prediction and prioritization Health records, medication records or vital signs A risk estimate, alert or ranked worklist for a care team Are the time horizon, threshold and response to an error defined?
Documentation and operations Speech, notes, messages or scheduling data A draft note, summary or work item for clinical or administrative staff Who checks omissions, controls access and signs the final record?
Patient-facing information Questions, forms or personal records A generated response or organized record for a patient or caregiver What is the stated scope, and when does the service transfer the interaction to a person?

This table is an APPI News editorial grouping, not a regulator's classification system. A scheduling assistant and software that supplies information for a clinical decision can use similar technology while carrying different evidence, governance and legal requirements.

Five stages connect data to a clinical workflow

A controlled workflow links the declared purpose, data, evaluation, human interaction and monitoring record. The five-stage sequence below is an APPI News synthesis of the cited WHO and IMDRF documents, not a certification program or a legal test in any country.

Stage one: Define the intended use

WHO's 2023 regulatory considerations place intended use, human intervention, continuous learning, cybersecurity, external validation and data quality inside the lifecycle for AI in health. A usable specification names the intended user, population, setting, input, output and decision that the output may influence.

The scope also records exclusions, missing-input behavior and the process that remains available when the system cannot produce a result. A broad phrase such as “assists clinicians” cannot determine a relevant test population, performance measure or response to failure.

Stage two: Build a traceable data pipeline

IMDRF's 10 good machine learning practice principles call for datasets that represent the intended population, use environment and measurement inputs, with training and test data kept appropriately independent and external validation proportionate to risk. The principles also identify subgroups and conditions in which a model may underperform as part of the evaluation.

A data record therefore needs provenance, collection dates and sites, equipment, units, missing-value handling, labeling methods and preprocessing. It also needs to connect each dataset to the model version that used it, because a later data pipeline or model release may no longer match the original test.

A clinician reviews AI analysis beside source data and model-version details (illustrative image)
Data provenance and model identity allow a result to be traced back to the system that produced it. (Illustrative image)

Stage three: Answer three evaluation questions

IMDRF's clinical-evaluation framework separates valid clinical association, analytical validation and clinical validation. Those tests ask whether the output has a sound relationship to the target condition, whether the software processes inputs reliably, and whether the output achieves its intended purpose in the target population and care context.

An overall accuracy score cannot answer all three questions or show what happens to false positives, false negatives and unusable inputs. A hospital validation record also needs independent data, subgroup results, workflow checks and defined responses when monitoring finds a problem.

Stage four: Connect the output to people and systems

The 2025 IMDRF principles call for assessing the human-AI team in the intended environment rather than evaluating the model alone. The assessment covers how users understand the output and its limits, possible overreliance, reasonably foreseeable misuse and the effects on work performance.

The operating design must state whether a person can accept, modify, reject or disregard the output, and it must preserve a fallback during an outage or unsupported case. An audit record can connect the input time, model and software version, output, reviewer and later action without treating the existence of a review button as proof of effective oversight.

Clinical and information teams review an AI output, workflow status and audit record (illustrative image)
Human authority, system integration and an audit trail are separate controls around the model. (Illustrative image)

Stage five: Monitor use and control changes

The IMDRF machine-learning principles call for risk-based monitoring in real-world use and controls for retraining risks such as overfitting, unintended bias and performance degradation. A monitoring record can set a baseline, review schedule, alert and stop conditions, investigation owner and recovery route for the deployed version.

The 2026 IMDRF framework divides a predetermined change control plan into a description of changes, a change plan and an impact assessment. It says the change plan should specify verification and validation, acceptance criteria, deployment and communication, while the impact assessment connects anticipated benefits, risks and mitigations.

The document is a common reference rather than a worldwide authorization, and it says some jurisdictions may not accept these plans for regulatory review. APPI News has separately examined how planned AI software changes require version-specific tests, user communication and a local hospital release record.

A hospital governance team reviews AI versions, monitoring signals and a stop procedure (illustrative image)
Monitoring needs a named response, and every update needs evidence tied to the released version. (Illustrative image)

One record keeps the five stages aligned

The workflow can be reviewed as five connected records, each describing the same intended use and deployed version. A separate APPI News analysis organizes deployment readiness into purpose, evidence, interoperability, human control and monitoring.

Record Minimum scope Mismatch it can expose
Purpose User, population, setting, input, output, affected decision, exclusions and fallback A marketing claim that is broader than the tested use
Data Provenance, sites, dates, equipment, preprocessing, labels and subgroup coverage Deployment data that differ from the evaluation sample
Evaluation Reference standard, independent test set, prespecified measures, errors and unresolved limits A headline score that does not support the intended decision
Workflow Interface, user authority, audit trail, training, outage handling and escalation A model result that reaches the wrong record or has no effective reviewer
Release and monitoring Version, baseline, alerts, stop authority, update evidence, notice and recovery A changed system that still relies on evidence for an earlier version

The records need common identifiers for the model, software, data pipeline and interface. If a supplier changes one component without updating the linked evidence, the file no longer describes the system in routine use.

The framework does not establish a worldwide verdict

WHO describes its 2023 publication as a high-level, nonexclusive overview rather than a regulatory framework or policy. IMDRF likewise publishes harmonized principles for jurisdictions to consider; its documents do not authorize a named product, assign liability or replace national and regional rules.

The good machine learning practice principles apply to AI-enabled medical devices, while the clinical-evaluation framework applies to software with a medical purpose. Administrative systems, research tools and patient-facing services may sit outside those scopes, although they still raise data, safety and governance questions.

Evidence also remains bound to the population, setting, inputs, workflow and version that were evaluated. Results from one hospital or dataset cannot establish performance after those conditions change.

APPI News could not find a comparable international dataset showing how many hospitals connect all five stages in routine medical AI use at the time of writing. Published principles describe the records organizations should consider, but they do not measure completion of the full workflow across countries.

A traceable medical AI workflow links a declared use to its data, evaluation, human interaction and deployed version. Monitoring then records whether the same system continues to operate within that evidence and what happened when a signal or software change crossed a preset boundary.