Detectable, Task-Relevant, or Harmful? Construct-Valid Evaluation of Unlabeled Distribution-Shift Monitoring in Chest Radiography

Distribution-shift alarms do not establish classifier harm. We evaluated the detectability, frozen-head relevance, and performance deterioration of two public DenseNet-121 chest-radiography classifiers under 30 controlled image conditions. Signals covered pixels, 64-component PCA representations, logits, probabilities, and calibrated high-error selection. An exact whole-model-null intervention changed representations while preserving all 18 released output slots to within 4.09 × 10−14; PCA monitoring detected all six interventions, five-logit monitoring detected none, and the paired Brier change was negligible. Across image shifts, PCA and logit alarm rates were 97% and 52%. In held-source/held-family prediction of primary Brier harm, probability mean displacement had the lowest error (MAE 0.00679); neither PCA signal outperformed the training-mean baseline after correction, whereas three output-facing signals beat both prespecified baselines. Secondary results were outcome-dependent: PCA RFF-MMD ranked best for macro-AUC and locked-F1 loss. An exploratory ConvNeXtV2-Atto/MobileViT-XS arm reproduced the construct separation: representation monitoring detected 57/60 conditions and all six task-null controls, while logit monitoring detected no task-null control; probability mean had the lowest held-architecture/held-family Brier-harm error ratio (0.404). The results show that detectability, task relevance, and metric-specific harm are distinct constructs and should be reported separately.

Authors

Institutions

Publication Details

Journal
Journal of Imaging
Published
2026-10-07
DOI
https://doi.org/10.3390/jimaging12100492
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Detectable, Task-Relevant, or Harmful? Construct-Valid Evaluation of Unlabeled Distribution-Shift Monitoring in Chest Radiography

Razvan Victor Rughinis, Dinu Țurcanu, Daniel Rosner, Dan Gabriel Badea et al.
Journal of Imaging
Explainable Artificial Intelligence (XAI)
article

Detectable, Task-Relevant, or Harmful? Construct-Valid Evaluation of Unlabeled Distribution-Shift Monitoring in Chest Radiography

Razvan Victor Rughinis, Dinu Țurcanu, Daniel Rosner, Dan Gabriel Badea, Flavia Zaim-Oprea, Răzvan-Andrei Rotaru
article en

Abstract

Distribution-shift alarms do not establish classifier harm. We evaluated the detectability, frozen-head relevance, and performance deterioration of two public DenseNet-121 chest-radiography classifiers under 30 controlled image conditions. Signals covered pixels, 64-component PCA representations, logits, probabilities, and calibrated high-error selection. An exact whole-model-null intervention changed representations while preserving all 18 released output slots to within 4.09 × 10−14; PCA monitoring detected all six interventions, five-logit monitoring detected none, and the paired Brier change was negligible. Across image shifts, PCA and logit alarm rates were 97% and 52%. In held-source/held-family prediction of primary Brier harm, probability mean displacement had the lowest error (MAE 0.00679); neither PCA signal outperformed the training-mean baseline after correction, whereas three output-facing signals beat both prespecified baselines. Secondary results were outcome-dependent: PCA RFF-MMD ranked best for macro-AUC and locked-F1 loss. An exploratory ConvNeXtV2-Atto/MobileViT-XS arm reproduced the construct separation: representation monitoring detected 57/60 conditions and all six task-null controls, while logit monitoring detected no task-null control; probability mean had the lowest held-architecture/held-family Brier-harm error ratio (0.404). The results show that detectability, task relevance, and metric-specific harm are distinct constructs and should be reported separately.

Journal of ImagingVol. 12(10)
Technical University of Moldova (MD), Academia Oamenilor de Știință din România (RO), Universitatea Națională de Știință și Tehnologie Politehnica București (RO)
Openalex Percentile: Top 12%
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.