INFORMER—Interpretability-Founded Monitoring of Medical Image Deep Learning Models: Application to Chest X-ray Pathologies

Deep learning has demonstrated strong performance in medical imaging. However, its limited interpretability remains a major barrier to clinical trust and safe deployment. This limitation is particularly relevant in multi-label classification, where quality control methods are still underdeveloped and commonly rely only on model outputs, without incorporating gradient-level information that may better reflect prediction reliability. In this study, we propose a quality control framework for multi-label medical image classification that improves both reliability and interpretability. The framework includes a graph-based class-distinctiveness method that analyzes saliency-derived information to identify unreliable predictions, as well as a retrieval-based extension that provides case-based explanations for flagged outputs. The proposed methods were evaluated on the CheXpert dataset and compared with established output-based quality control approaches. Robustness was assessed using bootstrapped test sets, and differences in ranking across bootstrap samples were analyzed using the Wilcoxon signed-rank test. The proposed framework outperformed baseline methods, achieving a higher mean F1 score (0.574 vs. 0.563), while the best-performing variant showed higher sensitivity (0.752 vs. 0.700). In bootstrapped analyses, it achieved better mean ranks than the baseline approaches. Averaged across input-image noise levels of 0.001-0.005, under IxG-based evaluation, the proposed framework showed improvements in bootstrapped F1 over the baseline, with the retrieval-based variant achieving a 21.92% improvement and the corresponding non-retrieval variant achieving a 13.53% improvement. Clinicians further evaluated the retrieved examples to determine their relevance for interpreting flagged predictions. These findings indicate that gradient-level and graph-based analysis can enhance the effectiveness, transparency, and clinical applicability of quality control in multi-label medical image classification.

Authors

Institutions

Publication Details

Journal
Journal of Imaging Informatics in Medicine
Published
2026-09-15
DOI
https://doi.org/10.1007/s10278-026-02257-8
Primary Topic
COVID-19 diagnosis using AI
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

INFORMER—Interpretability-Founded Monitoring of Medical Image Deep Learning Models: Application to Chest X-ray Pathologies

Alexander Poellinger, Mauricio Reyes, Dwarikanath Mahapatra, Shelley Zixin Shu et al.
Journal of Imaging Informatics in Medicine
COVID-19 diagnosis using AI
article

INFORMER—Interpretability-Founded Monitoring of Medical Image Deep Learning Models: Application to Chest X-ray Pathologies

Alexander Poellinger, Mauricio Reyes, Dwarikanath Mahapatra, Shelley Zixin Shu, Aurélie Pahud de Mortanges
article en

Abstract

Deep learning has demonstrated strong performance in medical imaging. However, its limited interpretability remains a major barrier to clinical trust and safe deployment. This limitation is particularly relevant in multi-label classification, where quality control methods are still underdeveloped and commonly rely only on model outputs, without incorporating gradient-level information that may better reflect prediction reliability. In this study, we propose a quality control framework for multi-label medical image classification that improves both reliability and interpretability. The framework includes a graph-based class-distinctiveness method that analyzes saliency-derived information to identify unreliable predictions, as well as a retrieval-based extension that provides case-based explanations for flagged outputs. The proposed methods were evaluated on the CheXpert dataset and compared with established output-based quality control approaches. Robustness was assessed using bootstrapped test sets, and differences in ranking across bootstrap samples were analyzed using the Wilcoxon signed-rank test. The proposed framework outperformed baseline methods, achieving a higher mean F1 score (0.574 vs. 0.563), while the best-performing variant showed higher sensitivity (0.752 vs. 0.700). In bootstrapped analyses, it achieved better mean ranks than the baseline approaches. Averaged across input-image noise levels of 0.001-0.005, under IxG-based evaluation, the proposed framework showed improvements in bootstrapped F1 over the baseline, with the retrieval-based variant achieving a 21.92% improvement and the corresponding non-retrieval variant achieving a 13.53% improvement. Clinicians further evaluated the retrieved examples to determine their relevance for interpreting flagged predictions. These findings indicate that gradient-level and graph-based analysis can enhance the effectiveness, transparency, and clinical applicability of quality control in multi-label medical image classification.

Journal of Imaging Informatics in Medicine
University of Bern (CH), Khalifa University of Science and Technology (AE), University Hospital of Bern (CH), Universidad Iberoamericana (DO)
National Science Foundation, University of Bern, Schweizerischer Nationalfonds zur Förderung der Wissenschaftlichen Forschung
Openalex Percentile: Top 12%
COVID-19 diagnosis using AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.