INFORMER—Interpretability-Founded Monitoring of Medical Image Deep Learning Models: Application to Chest X-ray Pathologies
Deep learning has demonstrated strong performance in medical imaging. However, its limited interpretability remains a major barrier to clinical trust and safe deployment. This limitation is particularly relevant in multi-label classification, where quality control methods are still underdeveloped and commonly rely only on model outputs, without incorporating gradient-level information that may better reflect prediction reliability. In this study, we propose a quality control framework for multi-label medical image classification that improves both reliability and interpretability. The framework includes a graph-based class-distinctiveness method that analyzes saliency-derived information to identify unreliable predictions, as well as a retrieval-based extension that provides case-based explanations for flagged outputs. The proposed methods were evaluated on the CheXpert dataset and compared with established output-based quality control approaches. Robustness was assessed using bootstrapped test sets, and differences in ranking across bootstrap samples were analyzed using the Wilcoxon signed-rank test. The proposed framework outperformed baseline methods, achieving a higher mean F1 score (0.574 vs. 0.563), while the best-performing variant showed higher sensitivity (0.752 vs. 0.700). In bootstrapped analyses, it achieved better mean ranks than the baseline approaches. Averaged across input-image noise levels of 0.001-0.005, under IxG-based evaluation, the proposed framework showed improvements in bootstrapped F1 over the baseline, with the retrieval-based variant achieving a 21.92% improvement and the corresponding non-retrieval variant achieving a 13.53% improvement. Clinicians further evaluated the retrieved examples to determine their relevance for interpreting flagged predictions. These findings indicate that gradient-level and graph-based analysis can enhance the effectiveness, transparency, and clinical applicability of quality control in multi-label medical image classification.
Authors
- Alexander Poellinger (ORCID: https://orcid.org/0000-0002-3998-0603)
- Mauricio Reyes (ORCID: https://orcid.org/0000-0002-2434-9990)
- Dwarikanath Mahapatra (ORCID: https://orcid.org/0000-0001-9749-7858)
- Shelley Zixin Shu
- Aurélie Pahud de Mortanges
Institutions
- University of Bern (CH)
- Khalifa University of Science and Technology (AE)
- University Hospital of Bern (CH)
- Universidad Iberoamericana (DO)
Publication Details
- Journal
- Journal of Imaging Informatics in Medicine
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1007/s10278-026-02257-8
- Primary Topic
- COVID-19 diagnosis using AI
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Science Foundation
- University of Bern
- Schweizerischer Nationalfonds zur Förderung der Wissenschaftlichen Forschung