Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification
Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ within a multi-task Deep Learning (DL) framework for MRI-based glioma diagnosis that performs tumor segmentation and predicts Isocitrate Dehydrogenase (IDH) mutation status, 1p/19q co-deletion status and tumor grade. We use Monte Carlo Dropout (MCD) as the primary sampling-based UQ method for the detailed task-aware analysis, obtaining predictive uncertainty and its aleatoric and epistemic components. We assess uncertainty along complementary axes: (i) MC sample convergence of uncertainty estimates and their decomposition into aleatoric and epistemic components, (ii) calibration of predictive probabilities, and (iii) operational utility of uncertainty estimates, including error detection, selective prediction, and associations with tumor segmentation performance. We also compare MCD with Deep Ensembles (DE) and Monte Carlo Deep Ensembles (MCDE) to assess whether the observed operational utility of uncertainty estimates extends beyond a single UQ method. For the tumor segmentation task, we further examine how different voxel-wise uncertainty aggregation strategies influence case-level reliability assessment, thereby explicitly accounting for task-specific uncertainty representation. Additionally, we study task interactions to quantify how tumor segmentation quality and uncertainty relate to the prediction of the tumor features. Finally, we explore whether a composite trust score integrating tumor segmentation and classification uncertainty improves error detection. Across tasks, uncertainty estimates supported meaningful error detection, while calibration depended on the dropout rate, with moderate rates yielding the most reliable predictive probabilities. Decomposition provided task-dependent interpretability but did not consistently improve error detection over predictive uncertainty alone. The comparison with DE and MCDE showed that ensemble-based uncertainty estimates provided comparable operational utility, although no UQ method consistently dominated across all tasks and metrics. Moreover, the proposed trust score did not consistently outperform classification uncertainty for selective prediction, indicating that task-specific predictive uncertainty remains the most informative operational indicator of trust. Overall, our results provide a task-aware evaluation strategy and practical guidance towards the development of trustworthy AI for glioma diagnosis.
Authors
- Carolin M. Pirkl (ORCID: https://orcid.org/0000-0002-5759-5290)
- Marion Smits (ORCID: https://orcid.org/0000-0001-5563-2871)
- Stefan Klein (ORCID: https://orcid.org/0000-0003-4449-6784)
- Sebastian R. van der Voort (ORCID: https://orcid.org/0000-0002-6526-8126)
- Sandeep S Kaushik (ORCID: https://orcid.org/0000-0003-0654-0799)
- Gonzalo Esteban Mosquera Rojas
Institutions
- Erasmus MC (NL)
- Erasmus MC Cancer Institute (NL)
- Amsterdam University Medical Centers (NL)
- Medical Delta (NL)
- University of Amsterdam (NL)
Publication Details
- Journal
- The Journal of Machine Learning for Biomedical Imaging
- Published
- 2026-09-28
- DOI
- https://doi.org/10.59275/j.melba.2026-456d
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- article
- Field-Weighted Citation Impact
- 0.00