| hide |
|
|---|
Clinical benchmarks are essential for evaluating new machine learning models in healthcare because they ensure relevance and impact. By setting a comparison standard grounded in real-world clinical practice, these benchmarks help establish whether a model provides tangible benefits over existing methods. This relevance ensures models address actual clinical needs rather than producing technically interesting but clinically irrelevant results.
Furthermore, benchmarks foster trust and safety. They align models with regulatory standards, reducing risks and proving the model's utility before deployment. Comparative benchmarking also helps stakeholders gauge the model’s effectiveness against current methods, ensuring that any advancements are not just statistically significant but clinically meaningful, thereby justifying the model’s use in healthcare settings.
Biases that should be considered during this phase of the project include sampling, convergence and participation bias. Additionally, using known societal inequities and biases that were identified during the literature search of the problem space, assess your retrospective dataset to ensure it is properly representative of these known issues. This can be achieved most simply by looking at dataset distributions, and minority and majority class representations. Furthermore, for protected or minority groups, a statistical power analysis should be performed to ensure there are enough individuals in the dataset to garner meaningful information.
Defining appropriate statistical measures of performance for an AI-solution is crucial to ensuring its effectiveness and reliability. The choice of metrics should align with the specific problem domain, objectives, and nature of the dataset. For instance, in binary classification tasks, measures like AUC-ROC (Area Under the Receiver Operating Characteristic Curve)36,37 and PR-AUC (Precision-Recall Area Under the Curve)38,39 are often used. However, selecting between these metrics depends on the data characteristics and known imbalances. AUC-ROC provides an overall view of the model’s ability to discriminate between classes, but it may not be as informative in cases of class imbalance. PR-AUC, on the other hand, focuses on the precision and recall trade-off, making it a better choice when the positive class is rare or when false positives and false negatives have differing consequences. Failure to select the right metric can result in a misleading evaluation of the model’s performance and potentially biased performances for the majority class.
Some metrics also have known dependencies that must be considered to ensure accurate interpretation of performance. For example, the Dice Similarity Coefficient (DICE), commonly used in medical image segmentation tasks, is influenced by the size or volume of the segmentation40,41. Larger segmentations may artificially inflate the DICE score, while smaller ones might unfairly penalize it. Understanding these dependencies is essential to avoid obscuring the true performance of the model. Practitioners should complement such metrics with additional measures or normalization techniques to account for these biases and provide a more holistic evaluation.
Another key aspect is the validation strategy employed to assess the AI model. Techniques like k-fold cross-validation help to mitigate overfitting and provide a robust estimate of the model’s generalizability42,43. Using too few folds might lead to high variance in performance estimates, while excessively large numbers of folds can increase computational costs without significant gains in reliability. An inadequate validation process can lead to over-optimistic results on retrospective data that fail to generalize to new prospective data, undermining the solution’s practical applicability.
Overfitting is another common pitfall that arises when statistical measures are not appropriately applied or understood13. Overfitting occurs when a model learns the training data too well, capturing noise instead of generalizable patterns. This often happens when performance metrics are optimized exclusively on the training data or when complex models are used without adequate regularization. For instance, reporting high accuracy on training data while neglecting poor performance on unseen data can give a false sense of success. Similarly, reliance on a single metric, such as accuracy, in imbalanced datasets can obscure significant deficiencies, such as a model’s inability to correctly identify minority class instances.
Misusing statistical measures—whether by selecting inappropriate metrics, inadequately validating models, or overfitting to training data—can lead to incorrect conclusions and poor real-world performance. By thoughtfully addressing these considerations, practitioners can create AI-solutions that are both accurate and trustworthy.
Equity objectives for a project are informed by current inequities related to the problem space, and will fall on a spectrum from maintaining current inequity levels to reducing them; an AI-solution should never make inequity levels worse. To obtain the identified equity objective, appropriate and complementary fairness metrics are required.
Selecting appropriate measures of fairness is imperative to the assessment of developed AI-solutions. A few example measures are highlighted below, but investigators are encouraged to do a review of the literature to see if any newer and more relevant measures have been developed. 44–47
Designed to analytically assess equality and equity issues in order to give insight into the nature of the models performance.
- Group Unawareness: Group unawareness means that a model does not use a sensitive variable during prediction48. For example, in a simple linear regression, the coefficient of the sensitive variable would be set to 0. Group Unawareness can be assessed using metrics such as SHAP, or by looking at the statistical significance of a univariable Ordinary Least Squares Regression that has been fit to predict the outcome of interest can be predicted using a single feature. Caution when using Group Unawareness is needed in scenarios where the sensitive variable may be highly correlated with a proxy variable.
- Statistical Parity/Demographic Parity: measurement of whether a model predicts positive outcome at equal rates for each segment of a subgroup.49
- Equal Opportunity: measures whether individuals from each segment of a protected class, who are eligible/qualified, have the same probability of receiving a positive outcome (e.g. being offered participation in a clinical trial). However, assessment of eligibility/qualification can be subjective, and Equal Opportunity does not address biases in the assessment process.46,50
- Equal Odds: this metric is used to quantify whether equal true positive rates and false positive rates exist between different groups. It is more restrictive than equal opportunity making it more appropriate when there are existing biases in the data. However, this is a very restrictive metric and may reduce the model's performance. 51,52
- Positive Predicted Value (PPV) - Parity: given a positive prediction, the precision is equal across different groups. For example, if a model positively predicts that a patient should receive “treatment X”, the probability of this treatment being successful is equal in all segments of a protected class. 53
- False Positive Rate (FPR) - Parity: this metric is the opposite of PPV-Parity and wants to ensure that each segment of a protected class has the same false positive rate. For example, if the model positively predicts that a patient should receive “treatment X”, the probability of this treatment not being successful is equal in all segments of the protected class. 54
- FairLearn: An open-source python package designed for metric calculation and reduction of bias in algorithms.
- Fairness Indicators: A package from Google designed to work with TensorFlow. Can be used for evaluation and visualization of group disparities.
- AIF360: An open-source package used to detect and mitigate biases. It is known for its extensive documentation and tutorials.
- Themis-ML: An open-source package for easy integration with scikit-learn. It is used mainly for binary classes and evaluation.
Assessment of an AI-model, or AI-solution, acquired from a third party for deployment in a healthcare institution is imperative for the safety of patients. Comprehensive questioning is required across multiple domains. Some general questions that should be asked include:
-
General Information
- What is the intended purpose and scope of the model?
- Has the model been used in clinical settings similar to ours?
- What are the specific clinical problems it aims to solve?
- Were patients and end-users consulted during the conception and development of this model?
-
Development and Validation 5. What data was used to train and validate the model? (e.g., size, sources, geographic diversity, demographic representation) Are we able to do an internal assessment of the data? 6. What validation processes were followed? (e.g., external validation, cross-validation) 7. What are the key performance metrics, and how do they vary across subpopulations of interest?
-
Bias and Fairness Assessment 8. Was the training data representative of the populations the model will serve? This may be challenging if the training data cannot be accessed and compared to local data. 9. What methods, if any, were used to detect and mitigate biases in the development phase? 10. Are there known performance disparities across demographic groups (e.g., age, sex, race, ethnicity)? 11. What fairness frameworks or metrics were used to evaluate the model? 12. Does the model include safeguards to minimize inequities in its recommendations? 13. Are there transparency mechanisms to report when biases or unfair outcomes are detected post-deployment?
-
Safety and Risk Management 14. What safeguards are in place to ensure patient safety in case of errors? 15. Has the model undergone stress testing for edge cases or rare clinical scenarios? 16. What adverse outcomes or unintended consequences were identified during testing? 17. Is there a protocol for monitoring and reporting errors or adverse events post-deployment?
-
Compliance and Regulatory Adherence 18. Does the model comply with relevant regulations? 19. Are there documented audit trails for data usage and model outputs?
-
Operational Integration 20. What are the technical requirements for integration with our existing systems (EHR, PACS, etc.)? 21. What resources are needed for deployment and maintenance? 22. What user training and support are provided?
-
Post-Deployment Monitoring 23. What tools or processes are available for continuous monitoring of model performance, clinical impact, and operational and regulatory adherence? 24. How frequently should the model be retrained or updated? 25. How is feedback from clinicians and patients incorporated into updates? 26. Are there mechanisms to adjust or recalibrate the model based on observed disparities?
-
Ownership and Intellectual Property 27. Who owns the model and any updates made during its use? 28. What are the terms of data sharing, if applicable?
When the retrospective dataset is split into training and testing cohorts (through e.g. randomized stratified splitting, bootstrapping, time wise split), investigators should take care to assure each cohort retains the same representation of known biases and inequities (i.e. the distributions should be the same when compared to the full retrospective dataset).
Data quality should also be assessed at this step. Data missingness and consistency in variable naming are two such tests that should be completed prior to commencing ML training. Additionally, if using longitudinal data, an understanding of whether there were any changes in clinical practice over time is required to avoid temporal bias (e.g. variable naming standards, standard treatment regimens, pandemics affecting patient presentation).
Similarly to the fairness and equity measures section presented in subsection 3.3 of our appendix, this section is used only to highlight a few different methods to improve model fairness. Additional methods can be found in some of the open-source fairness metric libraries mentioned above. A literature review is recommended if none of the methods below are appropriate, or to determine if there are newer and more relevant methods that have been developed.
In some scenarios, it is appropriate to change or adjust the dataset to be fairer before training an ML model.
- Relabeling and perturbation: These methods involve changing the end-point label or features in a dataset. Relabeling attempts to balance the dataset by changing the end-point label, while perturbation involves varying the features/variables to create a more balanced representation of the data. Two examples of these methods are disparate impact remover55 and “massaging”56. However, relabeling or perturbing data can introduce inaccuracies and distort the data's original distribution. This may lead to models learning incorrect or overly simplified relationships, so these techniques must be thoroughly validated against the original data distributions.
- Sampling: A dataset can be sampled up or down by adding or removing samples, respectively, to change the sample distributions to one that is more balanced. Up sampling the minority class can be done using duplication of existing samples or synthetic data, but caution is warranted since the model may start to overfit to duplicated or synthetic examples. Down sampling can involve removing majority group samples, but similarly, caution is warranted since data complexity can be reduced. Methods such as Synthetic Minority Over-sampling Technique (SMOTE)57 attempt to balance these methods.
These methods are used to modify or alter ML training algorithms to improve model fairness.
- Regularization and constraints: these methods alter the loss function of an algorithm. During regularization, an extra term is used to penalize discrimination, while constraints are used to limit the allowed bias level according to a certain loss function. Prejudice Remover58, Exponentiated Gradient Reduction59, Grid Search Reduction59 and Meta Fair Classifier60 are all examples of these techniques. A potential drawback of these methods is the difficulty in correctly defining the penalization terms or constraints. Overly strict constraints might lead to underfitting and reduced model performance, while poorly chosen terms can fail to adequately address bias. Additionally, these techniques may require significant computational resources and careful tuning, which can be challenging in practice.
- Adversarial learning involves training two models that compete to improve their performance61. Specifically, one model attempts to predict the true label of a dataset, while the other model attempts to exploit a known fairness issue using equality metrics. A drawback of adversarial learning is its complexity and the risk of instability during training, as the competing objectives of the models can lead to convergence issues. Adversarial models may inadvertently reduce overall predictive accuracy if fairness constraints conflict significantly with optimizing performance. Furthermore, adversarial models are not appropriate in scenarios where there are known differences between subgroups. As an example, when developing auto-segmentation models with a known performance difference between males and females adversarial learning should not be used since morphological differences are known to exist between sexes 62.
Post-processing methods act on the model predictions and are used when access to training data or the model is limited. These methods would be more appropriately used in scenarios where the model is commissioned from an external group.
- Classifier correction: A trained ML model is adapted to remove discrimination based on equalized odds and equality of opportunity constraints. Calibrated Equalized Odds63 is an example of one of these methods. However, classifier correction depends on accurately identifying sources of bias and setting appropriate constraints. Poorly defined constraints can lead to either insufficient fairness improvements or a decrease in the model’s predictive accuracy. Additionally, applying such corrections post-training may not address deeper issues of bias inherent in the data, limiting their effectiveness in mitigating unfairness.
- Output correction: model outputs are modified to obtain fairer distribution of the data. Reject Option based Classification64 is an example of this type of method that assigns more favourable outcomes to protected groups based upon low confidence regions of the classifier. While output correction can improve fairness metrics, it may distort predicted probabilities and reduce trust in model predictions. This approach addresses bias at the output level without resolving biases present in the training data or the model itself, potentially leading to superficial fairness improvements that fail to generalize to new datasets. Furthermore, aggressive corrections can harm overall model performance and usability for certain tasks.