A multidimensional trustworthiness evaluation framework for AI-enabled clinical decision support: a HER2 breast cancer case study

Abstract Machine learning models are increasingly being incorporated into AI-enabled decision-support systems, particularly in high-stakes domains such as healthcare. However, evaluation of such systems continues to rely predominantly on predictive performance metrics, which provide only a partial assessment of whether an AI-based decision-support system can be considered trustworthy for consequential decision-making. Recent advances in trustworthy artificial intelligence (AI) emphasize that characteristics such as explainability, reliability, robustness, and domain relevance should also be considered when evaluating AI systems intended to support decision-making. This study proposes a multidimensional trustworthiness evaluation framework for binary machine learning systems that integrates five dimensions: Predictive Reliability, Explainability, Explanation Stability, Clinical Plausibility, and Imbalance Robustness. The framework is designed as a reusable evaluation methodology rather than as a new classifier, feature-selection method, or explainability algorithm. Its empirical utility was examined through a retrospective case-study evaluation of HER2 status prediction in the METABRIC breast cancer dataset as a representative high-stakes clinical decision-support case. The evaluation included 1980 cases, comprising 1733 HER2-negative and 247 HER2-positive cases, and three classifiers—Decision Tree, Support Vector Machine (SVM), and XGBoost. Models were evaluated using 30 independent repetitions of stratified 80/20 holdout evaluation, with SHAP and LIME used for model explanation. Predictive performance was assessed using accuracy, precision, recall, F1-score, ROC-AUC, and PR-AUC, while trustworthiness dimensions were operationalized using ROC-AUC, SHAP–LIME agreement, explanation stability, clinical plausibility, and G-mean, respectively. XGBoost achieved the highest ROC-AUC (0.983 ± 0.008), PR-AUC (0.921 ± 0.029), F1-score (0.864 ± 0.024), explanation stability (0.652 ± 0.045), and clinical plausibility (0.303 ± 0.072). Decision Tree achieved the highest SHAP–LIME agreement (0.527 ± 0.105), whereas SVM achieved the highest G-mean (0.926 ± 0.023). Friedman tests identified significant differences among classifiers across all five trustworthiness dimensions ( p < 0.001 for each dimension). Pairwise analyses further demonstrated dimension-specific differences, including significant contrasts between all model pairs for predictive reliability, while Decision Tree and XGBoost did not differ significantly in explainability (Holm-adjusted p = 0.054) or clinical plausibility (Holm-adjusted p = 0.474). These findings demonstrate that model ranking can vary across trustworthiness dimensions and that predictive performance alone does not provide a sufficient characterization of AI systems intended to support consequential decisions. The framework therefore provides a structured multidimensional approach for evaluating binary machine learning systems, with potential relevance to trustworthy AI-enabled decision support beyond the clinical case examined here. The HER2 case study should not be interpreted as evidence of prospective clinical deployment readiness.

Authors

Institutions

Publication Details

Journal
Future Business Journal
Published
2026-09-28
DOI
https://doi.org/10.1186/s43093-026-01017-y
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A multidimensional trustworthiness evaluation framework for AI-enabled clinical decision support: a HER2 breast cancer case study

Hoda Waguih
Future Business Journal
Explainable Artificial Intelligence (XAI)
article

A multidimensional trustworthiness evaluation framework for AI-enabled clinical decision support: a HER2 breast cancer case study

Hoda Waguih
article en

Abstract

Abstract Machine learning models are increasingly being incorporated into AI-enabled decision-support systems, particularly in high-stakes domains such as healthcare. However, evaluation of such systems continues to rely predominantly on predictive performance metrics, which provide only a partial assessment of whether an AI-based decision-support system can be considered trustworthy for consequential decision-making. Recent advances in trustworthy artificial intelligence (AI) emphasize that characteristics such as explainability, reliability, robustness, and domain relevance should also be considered when evaluating AI systems intended to support decision-making. This study proposes a multidimensional trustworthiness evaluation framework for binary machine learning systems that integrates five dimensions: Predictive Reliability, Explainability, Explanation Stability, Clinical Plausibility, and Imbalance Robustness. The framework is designed as a reusable evaluation methodology rather than as a new classifier, feature-selection method, or explainability algorithm. Its empirical utility was examined through a retrospective case-study evaluation of HER2 status prediction in the METABRIC breast cancer dataset as a representative high-stakes clinical decision-support case. The evaluation included 1980 cases, comprising 1733 HER2-negative and 247 HER2-positive cases, and three classifiers—Decision Tree, Support Vector Machine (SVM), and XGBoost. Models were evaluated using 30 independent repetitions of stratified 80/20 holdout evaluation, with SHAP and LIME used for model explanation. Predictive performance was assessed using accuracy, precision, recall, F1-score, ROC-AUC, and PR-AUC, while trustworthiness dimensions were operationalized using ROC-AUC, SHAP–LIME agreement, explanation stability, clinical plausibility, and G-mean, respectively. XGBoost achieved the highest ROC-AUC (0.983 ± 0.008), PR-AUC (0.921 ± 0.029), F1-score (0.864 ± 0.024), explanation stability (0.652 ± 0.045), and clinical plausibility (0.303 ± 0.072). Decision Tree achieved the highest SHAP–LIME agreement (0.527 ± 0.105), whereas SVM achieved the highest G-mean (0.926 ± 0.023). Friedman tests identified significant differences among classifiers across all five trustworthiness dimensions ( p < 0.001 for each dimension). Pairwise analyses further demonstrated dimension-specific differences, including significant contrasts between all model pairs for predictive reliability, while Decision Tree and XGBoost did not differ significantly in explainability (Holm-adjusted p = 0.054) or clinical plausibility (Holm-adjusted p = 0.474). These findings demonstrate that model ranking can vary across trustworthiness dimensions and that predictive performance alone does not provide a sufficient characterization of AI systems intended to support consequential decisions. The framework therefore provides a structured multidimensional approach for evaluating binary machine learning systems, with potential relevance to trustworthy AI-enabled decision support beyond the clinical case examined here. The HER2 case study should not be interpreted as evidence of prospective clinical deployment readiness.

Future Business JournalVol. 12(1)
Sadat Academy for Management Sciences (EG)
Peace, Justice and strong institutions
Openalex Percentile: Top 9%
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.