A multidimensional trustworthiness evaluation framework for AI-enabled clinical decision support: a HER2 breast cancer case study
Abstract Machine learning models are increasingly being incorporated into AI-enabled decision-support systems, particularly in high-stakes domains such as healthcare. However, evaluation of such systems continues to rely predominantly on predictive performance metrics, which provide only a partial assessment of whether an AI-based decision-support system can be considered trustworthy for consequential decision-making. Recent advances in trustworthy artificial intelligence (AI) emphasize that characteristics such as explainability, reliability, robustness, and domain relevance should also be considered when evaluating AI systems intended to support decision-making. This study proposes a multidimensional trustworthiness evaluation framework for binary machine learning systems that integrates five dimensions: Predictive Reliability, Explainability, Explanation Stability, Clinical Plausibility, and Imbalance Robustness. The framework is designed as a reusable evaluation methodology rather than as a new classifier, feature-selection method, or explainability algorithm. Its empirical utility was examined through a retrospective case-study evaluation of HER2 status prediction in the METABRIC breast cancer dataset as a representative high-stakes clinical decision-support case. The evaluation included 1980 cases, comprising 1733 HER2-negative and 247 HER2-positive cases, and three classifiers—Decision Tree, Support Vector Machine (SVM), and XGBoost. Models were evaluated using 30 independent repetitions of stratified 80/20 holdout evaluation, with SHAP and LIME used for model explanation. Predictive performance was assessed using accuracy, precision, recall, F1-score, ROC-AUC, and PR-AUC, while trustworthiness dimensions were operationalized using ROC-AUC, SHAP–LIME agreement, explanation stability, clinical plausibility, and G-mean, respectively. XGBoost achieved the highest ROC-AUC (0.983 ± 0.008), PR-AUC (0.921 ± 0.029), F1-score (0.864 ± 0.024), explanation stability (0.652 ± 0.045), and clinical plausibility (0.303 ± 0.072). Decision Tree achieved the highest SHAP–LIME agreement (0.527 ± 0.105), whereas SVM achieved the highest G-mean (0.926 ± 0.023). Friedman tests identified significant differences among classifiers across all five trustworthiness dimensions ( p < 0.001 for each dimension). Pairwise analyses further demonstrated dimension-specific differences, including significant contrasts between all model pairs for predictive reliability, while Decision Tree and XGBoost did not differ significantly in explainability (Holm-adjusted p = 0.054) or clinical plausibility (Holm-adjusted p = 0.474). These findings demonstrate that model ranking can vary across trustworthiness dimensions and that predictive performance alone does not provide a sufficient characterization of AI systems intended to support consequential decisions. The framework therefore provides a structured multidimensional approach for evaluating binary machine learning systems, with potential relevance to trustworthy AI-enabled decision support beyond the clinical case examined here. The HER2 case study should not be interpreted as evidence of prospective clinical deployment readiness.
Authors
- Hoda Waguih (ORCID: https://orcid.org/0000-0001-8184-7948)
Institutions
- Sadat Academy for Management Sciences (EG)
Publication Details
- Journal
- Future Business Journal
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1186/s43093-026-01017-y
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- article
- Field-Weighted Citation Impact
- 0.00