The International Medical Device Regulators Forum (IMDRF) finalized 10 good machine learning practice principles for artificial intelligence (AI) in medical devices in January 2025. The document treats intended purpose, representative evidence, human-AI interaction and post-deployment monitoring as parts of one product lifecycle, leaving a model benchmark as only one piece of the deployment case.
Hospital teams can translate those principles into five implementation controls: purpose, evidence, interoperability, human control and monitoring after release. The sequence is an APPI News synthesis, not an IMDRF certification scheme, a regulatory approval or a legal test in any country.
Deployment begins with a decision, not an algorithm
A demonstration usually gives a model a prepared input and displays an output. Routine use adds patient identification, data retrieval, time pressure, incomplete records, staff interfaces, fallback procedures and a decision that may affect care.
A deployment specification therefore needs to identify the intended user, patient population, input, output and decision that the output may influence. It also needs to record when the system should not be used, what happens when it cannot produce a result and which person retains authority over the decision.
The World Health Organization's 2021 principles say humans should remain in control of health systems and medical decisions and call for accountability and redress when people are harmed by algorithm-based decisions. Those are governance requirements for the service around a model, not properties that can be established by its accuracy score.
Five controls turn a model into a clinical system
First control: Define the intended use before setting a test
A phrase such as “assists clinicians” is too broad to determine what evidence matters. A usable scope names the target population, intended users, input source and format, output, care setting, expected response time and exclusions. It also states whether the output prioritizes a worklist, flags a possible finding, drafts text or directly informs another type of decision.
Joint 2024 principles from the US Food and Drug Administration, Health Canada and the United Kingdom's Medicines and Healthcare products Regulatory Agency say transparency should cover a device's medical purpose, target population, intended users, use environment, inputs, outputs and intended effect on decisions. The document presents good practices rather than a law that applies worldwide.
The scope determines the acceptance test. A system intended to rank cases for later review needs a different response to failure than one that produces information during an active consultation. If that difference is absent from the specification, a passing test cannot show that the model is ready for its stated workflow.
Second control: Test representative evidence under clinical conditions
The IMDRF principles call for clinical evaluation with data that represent the intended population, use environment and measurement inputs; they also call for independent training and test data and external validation proportionate to risk. Testing should examine the device in clinically relevant conditions and assess the human-AI interaction rather than the model in isolation.
Representation must follow the stated use. Evidence from one patient mix, scanner fleet or record system cannot by itself establish how a model performs after those conditions change. A hospital evaluation can prespecify important subgroups, missing-data conditions, equipment variations and workflow interruptions, then report where the available sample is too small to support a conclusion.
A headline accuracy figure can also hide different error patterns. The review record needs the measures that match the decision, results for relevant subgroups and a clear account of false positives, false negatives and unusable inputs. A result outside the validated boundary should trigger the fallback defined in the first control.
Third control: Validate data exchange, meaning and access
Health Level Seven International (HL7) defines Fast Healthcare Interoperability Resources (FHIR) as a standard for exchanging health information electronically, with modular Resources that can be combined and linked for specific uses. The standard gives systems a common structure for concepts such as patients, observations, diagnostic reports and medications. It does not make every source record complete or give local codes the same meaning.
Interface testing has to go beyond receiving a successful application programming interface (API) response. Patient identifiers, units, code systems, timestamps, missing values, source context and version identifiers need expected results that both technical and clinical teams can inspect. The implementation also needs to show that an output reaches the correct user without stripping away information needed to interpret it.
HL7 explicitly says FHIR is not a security protocol and that authentication, authorization and access-control decisions require a separate security system. The specification supplies resources that can support provenance and audit records, but an organization still has to define who may retrieve each record and how access is reviewed. A separate APPI News report examines how Taiwan is building a FHIR layer for hospital records and medical AI.
Fourth control: Put human authority into the workflow
The label “human in the loop” does not describe what a person can do. A working design defines whether a clinician may accept, modify, reject or mark an output unusable, how the system presents uncertainty and what path remains available during an outage. It also sets a route for reporting an error instead of treating every override as user resistance.
The joint transparency principles focus on the performance of the human-AI team and say information should appear when it is needed in the workflow, including notifications about updates. That means a static user manual is not enough when a changed model affects what appears on screen. Interface tests need to show that status, limitations and warnings are visible at the decision point.
An audit record can connect the model version, input time, output and subsequent human action. That record helps an organization reconstruct an incident and compare performance across versions, but it also contains sensitive operational or patient information that needs access and retention controls. The responsible teams and escalation route have to be named before the system enters routine use.
Fifth control: Monitor performance and control changes
The IMDRF principles treat deployment as the start of another evidence phase. They call for monitoring performance in real-world use and managing retraining risks such as unintended bias, overfitting and degradation caused by data drift. A monitoring plan therefore needs a baseline, review schedule, alert thresholds, investigation process, stop condition and rollback path.
Final US Food and Drug Administration guidance issued in August 2025 recommends that a predetermined change control plan describe planned modifications, the methods for developing, validating and implementing them, and an assessment of their effects. The agency reviews such plans within US marketing submissions for devices in the 510(k), De Novo and premarket approval pathways.
The US document is nonbinding guidance for specified US pathways, not a worldwide change rule. Its structure can still expose missing information in a hospital contract: which changes are permitted, which evidence is required, who authorizes release and how a previous version can be restored. Regulatory status and change requirements must be checked separately in every market where a device is used.
The evidence package has to connect the five controls
A pre-deployment file can combine the intended-use specification, data description, independent evaluation, interface mapping, clinical workflow, audit design, monitoring plan and change-control process. Each document should point back to the same population, inputs, outputs and decision. A mismatch, such as a test population narrower than the procurement claim, is evidence that the controls are not yet aligned.
Procurement scoring can separate model evidence, integration readiness and operational governance. A high benchmark score cannot compensate for an interface that loses units, a workflow with no fallback or a contract that permits unreviewed model changes. Conversely, a technically sound interface does not establish clinical performance.
The five controls also create separate decision points. A hospital can stop after the evidence review without spending on integration, pause an interface that fails semantic checks or roll back a changed version without discarding the original procurement record. That separation makes it possible to identify which claim failed instead of treating the entire project as a single pass-or-fail demonstration.
What the framework cannot establish
The cited documents are international principles, technical specifications and guidance from named regulators. They do not classify a particular product, authorize it for sale, set one evidence threshold for every clinical use or assign liability after an error. Those questions depend on the product, intended purpose, contracts and law in each country.
The five-control sequence is also not a published measure of hospital adoption. APPI News could not find a comparable international dataset showing how many hospitals use every control in routine medical AI deployments at the time of writing. Public documents often describe a framework, pilot or regulatory pathway without reporting whether a specific hospital has completed clinical validation and post-release monitoring.
Together, the cited documents define the shape of the evidence problem: intended purpose sets the boundary, representative testing checks performance, interoperability places the model in a real system, human control defines its role and monitoring follows changes after release. Within this five-control framework, the record is complete only when those documents describe the same system and the same clinical use.
Sources and further reading
- Good machine learning practice for medical device development: Guiding principles(International Medical Device Regulators Forum)
- Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles(US Food and Drug Administration, Health Canada and the United Kingdom's Medicines and Healthcare products Regulatory Agency)
- FHIR Overview(HL7 International)
- FHIR Security(HL7 International)
- WHO issues first global report on artificial intelligence in health(World Health Organization)
- Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions(US Food and Drug Administration)