Risk management for machine learning in medical devices
- 01ISO/TS 24971-2:2026 is a Technical Specification used with ISO 14971:2019 and ISO/TR 24971. It introduces no new process and no new requirements.
- 02It directs the process already in place at the causes of harm that arise when behaviour depends on data, statistical learning, and deployment context.
- 03Its scope is the machine learning enabled medical device. It excludes devices that use large language models or generative AI, so a governance claim built on this document alone does not cover generative behaviour.
- 04Clause 4.4 carries the one genuinely new requirement: the risk management plan must decide, during planning, whether post-production activities include performance monitoring, updates, or retraining.
- 05Clauses 4.2, 4.5, 6, 7.3 to 7.6, and 9 gain no MLMD-specific guidance. ISO/TR 24971 continues to apply.
What it is
ISO/TS 24971-2:2026 was published in June 2026. Rather than replacing the risk management process of ISO 14971 or adding requirements to it, it directs that process towards the sources of harm that exist because a device learns from data. Its scope is the machine learning enabled medical device, referred to throughout as an MLMD.
One post-market signal may challenge a model, deployment context, risk control, and release assumption simultaneously.
A Life-Cycle Approach to AI Governance in Medical Device Development, 2026
Two definitions of risk, and they cannot share a record
ISO/IEC 23894 and ISO 31000 define risk as the effect of uncertainty on objectives. ISO 14971 defines it as the combination of the probability of occurrence of harm and its severity. These perspectives overlap whenever an AI characteristic can contribute to harm, but neither contains the other, so neither can share a record with the other.
Set acceptance criteria first
Establish MLMD acceptance criteria during concept development, before the ML model is tested. Criteria written afterwards are shaped by the results they are meant to judge. They are not criteria for risk acceptability, which remain a separate manufacturer decision.
Seven places where machine learning changes the work
Clauses 4 to 7 of ISO 14971 do not change. What follows is the delta against the risk management file already in place. An analysis that stops at "incorrect output" as a software failure never asks why statistical performance varies across populations, acquisition devices, sites, or time, and is too coarse to select or monitor the most effective control.
| Clause | What it adds for an MLMD |
|---|---|
| 4.3 Competence | The team shall cover the ML life cycle: good machine learning practice, platform IT and security, algorithm and simulation testing, software validation of the risk control measures, usability engineering to IEC 62366-1, clinical workflow, and data management. |
| 4.4 Risk management plan | The one new requirement. Decide during planning whether post-production activities include performance monitoring, updates, or retraining. If monitoring is necessary for safety, the plan shall define the methods and processes and the data collected and reviewed. If retraining manages risk, the plan states the criteria for initiating it, how often they are applied, the retraining activities, and the provisions for returning to a previous version. |
| 5.2 Use and misuse | Reasonably foreseeable misuse includes use outside the intended patient population, use with inadequate input data, use error caused by limited transparency, and overreliance on the output. Where performance differs between patient groups, restricting the intended use is permitted. |
| 5.3 and 5.4 Hazards | Annex C adds MLMD-specific questions for identifying characteristics related to safety. Annex B adds examples of hazards, sequences of events, and hazardous situations. Use both with the questions in ISO/TR 24971, Annex A. |
| 5.5 Risk estimation | Where the probability of occurrence of harm cannot be estimated, estimate the risk on the severity of possible harm alone. Hardware and IT reliability complicate the estimate. Usability evaluation and quality assessment of the training and test data support a defensible estimate. |
| 7.1 Control options | Same order of priority. Inherently safe design first: choice of ML algorithm and model, data quality metrics, training and test data verified for completeness, correctness and consistency, evidence that test data were sequestered, input fields that accept only realistic data, and controlled access to data. Then protective measures: human oversight with an intervention mechanism, cross validation of the output, alarm signals, and measures for loss of connectivity. Then information for safety and the hand-off strategy. |
| 7.2 Verification | Test the trained model against the acceptance criteria set during concept development. Synthetic data can verify a risk control measure under simulated use conditions. Test a demographic bias control by resubmitting a patient record with one attribute changed. Verify return to a previous version by simulation, and confirm that no new risk is introduced and no existing risk increased. |
What to weigh at release, and what to watch after it
Clause 8.1, overall residual risk
Four factors are added to the evaluation. Silent failure: assume the MLMD can fail without informing the user, and include that case. Level of autonomy: autonomy can improve patient management and can reduce the situational awareness of the user. Overreliance: evaluate use-related residual risk for unwanted bias, including transparency, understandability, and over-trust. Novelty: where the MLMD changes current medical practice, analyze that degree of novelty. Record the acceptable range of MLMD performance here, because it is the baseline for post-production monitoring.
Clause 8.2, what to disclose
Disclose significant residual risks as before, with two additions, both written for the intended user. Transparency: performance specifications and limitations, including differences in performance between patient groups that result from the selection of training data, and the checks recommended before use. Explainability: the features that the algorithm uses and weights when the model parameters are determined, where this can be stated, and a full justification where it cannot.
Annex A, bias
Bias is a systematic difference in how objects, people, or groups are treated. It arises in the data, in the model, and in how a user reads the output. Each source needs a risk control measure and verification of its effectiveness. Dataset size is not a risk control measure.
Clause 10, production and post-production
Monitoring and change review are directed at drift, retraining, site differences, feedback loops, and rollback. Actively collect safety-related information, considering the level of autonomy, continuous learning, and how the device is used in clinical practice. Review for drift against the performance range recorded at release. Apply the criteria defined in the risk management plan: improve the concept, retrain the model, update the device, or return to a previous version. Record the outcome in the risk management file. Retraining and return to a previous version are changes, and remain subject to ISO 13485 controls.
The records have to stay connected
Every clause above assumes a record that survives change, because a reviewer has to be able to say which model version is released, which datasets trained and tested it, and which approval a new post-production signal invalidates. ISO/TS 24971-2 connects ISO 14971 to the causal factors specific to machine learning, but not how those dependencies stay retrievable after several retraining cycles.
A document repository can contain all of this information, but only a controlled relationship model makes the dependency query reliable. That means the AI System, AI Use Case, AI Model, Dataset, Dataset-Model Association, and AI Risk are distinct records, so a change to one shows which decisions it reopens. It means the two risk layers stay separate: ISO/IEC 23894 informs the wider AI risk method while ISO/TS 24971-2 directs ML-specific causal factors into the ISO 14971 analysis, without merging criteria. And it means boundaries, methods, evidence, and decision authority are defined before change is attempted, so a Predetermined Change Control Plan can be assembled from those records.
Six questions to ask about one deployed device
More than two answers of no indicates that the risk management file describes a device that is no longer released.
Summary and interpretation only. Qity supports manufacturers in applying ISO/TS 24971-2. The manufacturer retains responsibility for the risk management file.
A life-cycle approach to AI governance in medical device development
Miguel Azevedo, Qity, 2026 · 25 pages, open access
Qity runs free structured assessments at qity.app, covering which regulations apply to your company and whether a given product is ready to declare.



