Image Quality and Model Confidence for Failure Detection in Chest X-Ray Classification
This study investigates the reliability of deep learning-based chest X-ray classification under real-world image quality degradation. While classification models can achieve high accuracy on clean medical images, their predictions may become unreliable when images are affected by blur, noise, or other quality distortions. We evaluate both predictive performance and model confidence across clean and degraded image conditions, examining whether confidence scores can serve as useful signals for identifying potentially unreliable predictions. The study further analyzes calibration using Brier scores and Expected Calibration Error (ECE), alongside classification and confidence-based metrics. Results demonstrate that image degradation can affect both predictive reliability and confidence behavior, highlighting the importance of evaluating model confidence and calibration rather than relying solely on classification accuracy. The findings support the development of more reliable medical AI systems capable of recognizing when their predictions may be uncertain or potentially erroneous.
Authors
- Mohsin Raza Dahri
Institutions
- Quaid-e-Awam University of Engineering, Science and Technology (PK)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23148537
- Primary Topic
- COVID-19 diagnosis using AI
- Type
- preprint