"Pathology-Aware Explainability in Chest X-ray Classifiers

Evaluating chest radiograph classifiers presents a fundamental mismatch: models are conventionally validated by discrimination metrics, yet radiological diagnosis relies on the structured interpretation of visual signs—its semiology. This mismatch allows strong classification despite attribution to regions with no finding-specific meaning. We propose a semiology-grounded framework to evaluate the anatomical faithfulness of post-hoc explanations beyond generic geometric overlap. Using VinDr-CXR filtered for strict multi-radiologist consensus, we trained AlexNet and DenseNet-121 at two input resolutions and paired two findings with divergent spatial signatures—aortic enlargement and cardiomegaly—to distinguish genuine localisation from a central default. Grad-CAM attributions were cross-validated with gradient-free occlusion sensitivity and calibrated against an input-blind central-Gaussian baseline. Although all models achieved strong discrimination (Matthews Correlation Coefficient 0.56–0.74, AUC-ROC > 0.91), their explanations localised no better than the central baseline and produced highly overlapping activation across findings. Step-wise analysis ruled out the tested background shortcut and, through agreement with occlusion sensitivity, identified a model-intrinsic positional bias invisible to performance metrics. Anatomical supervision relocated AlexNet’s attributions onto the findings without sacrificing classification performance, whereas DenseNet-121 overfitted the same supervision. Controlled ablations indicate that correction depends on model capacity relative to available anatomical supervision. Notably, a penalty carrying no anatomical information reached almost the same localisation score as explicit box supervision, showing that localisation metrics, like discrimination metrics, can be satisfied without anatomical grounding. These findings show that neither classification nor localisation scores alone can establish anatomically faithful model evidence, which must instead be evaluated against finding-specific semiology and spatially informed baselines.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-05
DOI
https://doi.org/10.5281/zenodo.22336710
Primary Topic
COVID-19 diagnosis using AI
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

"Pathology-Aware Explainability in Chest X-ray Classifiers

Fernando Carlos López Hernández, Sonia Rubio Herranz, Antonio López Montes, Giorgio Venturini García
Zenodo (CERN European Organization for Nuclear Research)
COVID-19 diagnosis using AI
article

"Pathology-Aware Explainability in Chest X-ray Classifiers

Fernando Carlos López Hernández, Sonia Rubio Herranz, Antonio López Montes, Giorgio Venturini García
article en

Abstract

Evaluating chest radiograph classifiers presents a fundamental mismatch: models are conventionally validated by discrimination metrics, yet radiological diagnosis relies on the structured interpretation of visual signs—its semiology. This mismatch allows strong classification despite attribution to regions with no finding-specific meaning. We propose a semiology-grounded framework to evaluate the anatomical faithfulness of post-hoc explanations beyond generic geometric overlap. Using VinDr-CXR filtered for strict multi-radiologist consensus, we trained AlexNet and DenseNet-121 at two input resolutions and paired two findings with divergent spatial signatures—aortic enlargement and cardiomegaly—to distinguish genuine localisation from a central default. Grad-CAM attributions were cross-validated with gradient-free occlusion sensitivity and calibrated against an input-blind central-Gaussian baseline. Although all models achieved strong discrimination (Matthews Correlation Coefficient 0.56–0.74, AUC-ROC > 0.91), their explanations localised no better than the central baseline and produced highly overlapping activation across findings. Step-wise analysis ruled out the tested background shortcut and, through agreement with occlusion sensitivity, identified a model-intrinsic positional bias invisible to performance metrics. Anatomical supervision relocated AlexNet’s attributions onto the findings without sacrificing classification performance, whereas DenseNet-121 overfitted the same supervision. Controlled ablations indicate that correction depends on model capacity relative to available anatomical supervision. Notably, a penalty carrying no anatomical information reached almost the same localisation score as explicit box supervision, showing that localisation metrics, like discrimination metrics, can be satisfied without anatomical grounding. These findings show that neither classification nor localisation scores alone can establish anatomically faithful model evidence, which must instead be evaluated against finding-specific semiology and spatially informed baselines.

Zenodo (CERN European Organization for Nuclear Research)
Universidad Complutense de Madrid (ES), Department of Mathematical Sciences (RU)
Reduced inequalities
Openalex Percentile: Top 10%
COVID-19 diagnosis using AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.