Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification

Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ within a multi-task Deep Learning (DL) framework for MRI-based glioma diagnosis that performs tumor segmentation and predicts Isocitrate Dehydrogenase (IDH) mutation status, 1p/19q co-deletion status and tumor grade. We use Monte Carlo Dropout (MCD) as the primary sampling-based UQ method for the detailed task-aware analysis, obtaining predictive uncertainty and its aleatoric and epistemic components. We assess uncertainty along complementary axes: (i) MC sample convergence of uncertainty estimates and their decomposition into aleatoric and epistemic components, (ii) calibration of predictive probabilities, and (iii) operational utility of uncertainty estimates, including error detection, selective prediction, and associations with tumor segmentation performance. We also compare MCD with Deep Ensembles (DE) and Monte Carlo Deep Ensembles (MCDE) to assess whether the observed operational utility of uncertainty estimates extends beyond a single UQ method. For the tumor segmentation task, we further examine how different voxel-wise uncertainty aggregation strategies influence case-level reliability assessment, thereby explicitly accounting for task-specific uncertainty representation. Additionally, we study task interactions to quantify how tumor segmentation quality and uncertainty relate to the prediction of the tumor features. Finally, we explore whether a composite trust score integrating tumor segmentation and classification uncertainty improves error detection. Across tasks, uncertainty estimates supported meaningful error detection, while calibration depended on the dropout rate, with moderate rates yielding the most reliable predictive probabilities. Decomposition provided task-dependent interpretability but did not consistently improve error detection over predictive uncertainty alone. The comparison with DE and MCDE showed that ensemble-based uncertainty estimates provided comparable operational utility, although no UQ method consistently dominated across all tasks and metrics. Moreover, the proposed trust score did not consistently outperform classification uncertainty for selective prediction, indicating that task-specific predictive uncertainty remains the most informative operational indicator of trust. Overall, our results provide a task-aware evaluation strategy and practical guidance towards the development of trustworthy AI for glioma diagnosis.

Authors

Institutions

Publication Details

Journal
The Journal of Machine Learning for Biomedical Imaging
Published
2026-09-28
DOI
https://doi.org/10.59275/j.melba.2026-456d
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification

Carolin M. Pirkl, Marion Smits, Stefan Klein, Sebastian R. van der Voort et al.
The Journal of Machine Learning for Biomedical Imaging
Explainable Artificial Intelligence (XAI)
article

Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification

Carolin M. Pirkl, Marion Smits, Stefan Klein, Sebastian R. van der Voort, Sandeep S Kaushik, Gonzalo Esteban Mosquera Rojas
article en

Abstract

Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ within a multi-task Deep Learning (DL) framework for MRI-based glioma diagnosis that performs tumor segmentation and predicts Isocitrate Dehydrogenase (IDH) mutation status, 1p/19q co-deletion status and tumor grade. We use Monte Carlo Dropout (MCD) as the primary sampling-based UQ method for the detailed task-aware analysis, obtaining predictive uncertainty and its aleatoric and epistemic components. We assess uncertainty along complementary axes: (i) MC sample convergence of uncertainty estimates and their decomposition into aleatoric and epistemic components, (ii) calibration of predictive probabilities, and (iii) operational utility of uncertainty estimates, including error detection, selective prediction, and associations with tumor segmentation performance. We also compare MCD with Deep Ensembles (DE) and Monte Carlo Deep Ensembles (MCDE) to assess whether the observed operational utility of uncertainty estimates extends beyond a single UQ method. For the tumor segmentation task, we further examine how different voxel-wise uncertainty aggregation strategies influence case-level reliability assessment, thereby explicitly accounting for task-specific uncertainty representation. Additionally, we study task interactions to quantify how tumor segmentation quality and uncertainty relate to the prediction of the tumor features. Finally, we explore whether a composite trust score integrating tumor segmentation and classification uncertainty improves error detection. Across tasks, uncertainty estimates supported meaningful error detection, while calibration depended on the dropout rate, with moderate rates yielding the most reliable predictive probabilities. Decomposition provided task-dependent interpretability but did not consistently improve error detection over predictive uncertainty alone. The comparison with DE and MCDE showed that ensemble-based uncertainty estimates provided comparable operational utility, although no UQ method consistently dominated across all tasks and metrics. Moreover, the proposed trust score did not consistently outperform classification uncertainty for selective prediction, indicating that task-specific predictive uncertainty remains the most informative operational indicator of trust. Overall, our results provide a task-aware evaluation strategy and practical guidance towards the development of trustworthy AI for glioma diagnosis.

The Journal of Machine Learning for Biomedical ImagingVol. 2026(UNSURE2025)
Erasmus MC (NL), Erasmus MC Cancer Institute (NL), Amsterdam University Medical Centers (NL), Medical Delta (NL), University of Amsterdam (NL)
Openalex Percentile: Top 9%
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.