Image Quality and Model Confidence for Failure Detection in Chest X-Ray Classification

This study investigates the reliability of deep learning-based chest X-ray classification under real-world image quality degradation. While classification models can achieve high accuracy on clean medical images, their predictions may become unreliable when images are affected by blur, noise, or other quality distortions. We evaluate both predictive performance and model confidence across clean and degraded image conditions, examining whether confidence scores can serve as useful signals for identifying potentially unreliable predictions. The study further analyzes calibration using Brier scores and Expected Calibration Error (ECE), alongside classification and confidence-based metrics. Results demonstrate that image degradation can affect both predictive reliability and confidence behavior, highlighting the importance of evaluating model confidence and calibration rather than relying solely on classification accuracy. The findings support the development of more reliable medical AI systems capable of recognizing when their predictions may be uncertain or potentially erroneous.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23148537
Primary Topic
COVID-19 diagnosis using AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Image Quality and Model Confidence for Failure Detection in Chest X-Ray Classification

Mohsin Raza Dahri
Zenodo (CERN European Organization for Nuclear Research)
COVID-19 diagnosis using AI
preprint

Image Quality and Model Confidence for Failure Detection in Chest X-Ray Classification

Mohsin Raza Dahri
preprint en

Abstract

This study investigates the reliability of deep learning-based chest X-ray classification under real-world image quality degradation. While classification models can achieve high accuracy on clean medical images, their predictions may become unreliable when images are affected by blur, noise, or other quality distortions. We evaluate both predictive performance and model confidence across clean and degraded image conditions, examining whether confidence scores can serve as useful signals for identifying potentially unreliable predictions. The study further analyzes calibration using Brier scores and Expected Calibration Error (ECE), alongside classification and confidence-based metrics. Results demonstrate that image degradation can affect both predictive reliability and confidence behavior, highlighting the importance of evaluating model confidence and calibration rather than relying solely on classification accuracy. The findings support the development of more reliable medical AI systems capable of recognizing when their predictions may be uncertain or potentially erroneous.

Zenodo (CERN European Organization for Nuclear Research)
Quaid-e-Awam University of Engineering, Science and Technology (PK)
COVID-19 diagnosis using AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Image Quality and Model Confidence for Failure Detection in Chest X-Ray Classification — Mohsin Raza Dahri · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS