The International Medical Device Regulators Forum (IMDRF) published 10 good machine learning practice principles for artificial intelligence (AI) in medical devices in January 2025. The principles call for representative datasets, independent training and test data, clinically relevant testing, subgroup disclosure and monitoring after deployment.
A hospital can turn that lifecycle framework into an acceptance test that asks who the model serves, where its evidence came from and what each kind of error costs in the intended workflow. One overall accuracy score cannot answer those questions because it can conceal differences among patient groups, clinical sites, measurement devices and data-collection conditions.
The evidence base is uneven before a model reaches a hospital
A 2022 PLOS Digital Health review retrieved 30,576 records in a search of 2019 publications and deemed 7,314 eligible; in its pooled labeled samples, 40.8 percent of data sources came from the United States and 13.7 percent from China. Radiology accounted for 40.4 percent of the pooled studies, while pathology ranked second at 9.1 percent.
Those figures describe imbalances in published research and data sources, not proof that any individual model treats patients unfairly. They show why a paper's average result cannot establish that a model will generalize to a different population, hospital or clinical specialty.
The World Health Organization warned in 2021 that systems trained chiefly on data from high-income countries may not perform well in low- and middle-income settings and said health AI should reflect diverse socioeconomic and health care settings. Its six principles also call for inclusiveness, equity, accountability and continuous assessment during actual use.
Start with the intended use and a data map
For an acceptance test, bias is a recurring performance difference associated with the population, data source, measurement process or use environment. It can enter through missing patient groups, inconsistent labels, equipment differences, referral patterns, disease severity or the way staff act on an output.
The test specification should first name the intended patients, users, setting, inputs, outputs and decision that the output may influence. A data map can then record age, sex, gender, race, ethnicity, geography and medical condition where relevant, together with site type, equipment, acquisition protocol, care setting, collection period and missing-data patterns.
Each characteristic needs a count in the training, test and monitoring datasets. The record should mark groups that are absent, poorly represented or too small for a stable estimate instead of merging them silently into an aggregate score.
Subgroup definitions must fit the intended population and the data that local health systems can lawfully and consistently collect. Labels used in one country may not map cleanly to another, and a dataset described as international can still omit the sites, devices or patient groups relevant to a particular deployment.
Keep development data out of the acceptance test
Training and test data need enough separation to prevent the model from being evaluated on cases it has effectively seen before. Patient overlap is the most direct problem, but dependence can also arise when repeated examinations, related records, one site's acquisition process or one collection period appears on both sides of the split.
The IMDRF principles say teams should consider patient, site and data-acquisition dependencies and make external validation proportionate to risk. They do not set a universal split ratio, number of hospitals or minimum sample size, so the evaluation plan has to justify why its independent sample is adequate for the stated use and relevant subgroups.
The reference standard also needs its own record. That record should name how the correct outcome was established, who resolved disagreements, which limitations remain and whether the process was independent of the model output.
Report subgroup performance with uncertainty
The measures must match the model's task and the clinical cost of an error. A classification report may include sensitivity, specificity, positive and negative predictive values, false positives and false negatives, while a risk model may also need calibration across its operating range.
Joint 2024 principles from Health Canada, the US Food and Drug Administration and the United Kingdom's Medicines and Healthcare products Regulatory Agency say device information should cover known biases or failure modes, confidence intervals, poorly represented populations, input mismatches, local acceptance testing and ongoing monitoring. The document presents good practices for transparency, not a worldwide approval standard.
Each subgroup row should pair the selected measure with its sample count, numerator and denominator, operating threshold and confidence interval. A five-percentage-point difference based on a small sample may be unstable, while a test that finds no difference may simply lack the power to detect one.
Testing age and sex separately can also miss a problem concentrated in an age-and-sex group at one site or on one device. The evaluation should prespecify combinations tied to the intended use, but it should label sparse results as unresolved rather than treating a noisy ranking as a fairness verdict.
Test the human-AI system under clinical conditions
A silent retrospective test can estimate model performance without showing what happens when staff receive the output. Acceptance testing also needs to examine whether users understand the display, recognize excluded inputs, respond to unavailable results and avoid relying on the system outside its validated scope.
A simulated workflow or limited prospective evaluation can record output availability, review and override rates, time to action, unreviewed alerts and the workload created by false positives. It should also test missing fields, unsupported devices, delayed data, interface failures and other foreseeable conditions that change the input or prevent a result.
The failure cost determines which measures receive priority. A tool that prioritizes a worklist may need close review of false negatives and delayed cases, while a reminder system may create most of its burden through false positives; neither can be judged from the same aggregate threshold alone.
Monitoring must preserve subgroup and version detail
Passing one evaluation establishes performance only for the tested model version and conditions. Patient mix, equipment, referral pathways, input fields and user behavior can change, while retraining or a software update can alter the model itself.
The monitoring plan should retain the same subgroup definitions and task-specific measures used at acceptance, along with the deployed model version and input conditions. Where subgroup cases arrive slowly, the report can use a longer observation window but must state the delay and resulting uncertainty.
Investigation, restriction and suspension triggers need to be assigned before routine use begins. The operating record should identify who reviews a signal, which fallback remains available, how an earlier version can be restored and what evidence is required before the system returns to service.
A version change does not inherit the previous version's subgroup evidence automatically. Hospitals need to assess whether the model, interface, data pipeline, intended population or workflow changed and repeat the affected parts of the test.
An acceptance record needs six linked decisions
A hospital can organize the evidence into six records before routine use. This structure is an APPI News synthesis of the cited guidance, not a certification scheme or regulatory approval route, and it does not determine compliance with any country's requirements.
- Intended use: Name the patients, users, setting, inputs, output, decision influenced, exclusions and fallback.
- Data coverage: Record the origin, period, patient and site mix, equipment, missingness and representation of relevant groups.
- Independent evaluation: Document separation from development data, external sites, reference standards and unresolved dependencies.
- Subgroup results: Report task-specific measures, thresholds, sample counts, confidence intervals and evidence gaps beside the overall result.
- Workflow failures: Test the human-AI team, unsupported inputs, unavailable outputs, overrides, false-positive burden and fallback operation.
- Monitoring and change control: Link each result to a model version, review schedule, action threshold, responsible owner, suspension route and retest rule.
A pass means that the available evidence supports a defined use under defined conditions. It does not establish that the model is unbiased in every population, compatible with every hospital or accepted by regulators in every country.
APPI News could not find a published global dataset showing how many hospitals routinely report subgroup performance after deploying medical AI. The absence of that count limits comparisons of adoption, but it does not change the evidence each hospital needs to preserve for its own acceptance and monitoring decisions.
Sources and further reading
- Good machine learning practice for medical device development: Guiding principles(International Medical Device Regulators Forum)Final document, January 2025
- Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles(Health Canada, US Food and Drug Administration, and the United Kingdom's Medicines and Healthcare products Regulatory Agency)June 2024
- WHO issues first global report on artificial intelligence in health and six guiding principles for its design and use(World Health Organization)June 28, 2021
- Sources of bias in artificial intelligence that perpetuate healthcare disparities—A global review(PLOS Digital Health)March 31, 2022