Similarity‐Aware Evaluation, Probability Calibration, and Independent Replication for Reliable Gastrointestinal Image Classification

ABSTRACT High accuracy alone does not establish reliable medical‐image classification because visually similar images may cross partitions and confidence may be miscalibrated. We evaluated a similarity‐aware framework on a four‐class gastrointestinal benchmark (4000 images) and independently replicated the methodology on GastroVision (8000 images; 27 classes). A 64‐bit DCT perceptual hash defined groups at Hamming distance ≤ 4. Five‐fold random and group‐constrained protocols were matched by model and seed, with separate validation, calibration, and test subsets. We compared modern backbones, temperature scaling, a 2 × 2 Mixup/label‐smoothing ablation, and deterministic maximum‐softmax probability. On the primary benchmark, grouped EfficientNet‐B0 achieved accuracy and macro‐F1 of 0.9765 ± 0.0022. Grouping changed accuracy by +0.0025 relative to random splitting (Wilcoxon p = 0.625). Temperature scaling reduced expected calibration error from 0.1292 ± 0.0216 to 0.0123 ± 0.0030 without changing predictions, and deterministic MSP achieved error‐detection AUROC of 0.9288 ± 0.0181. On GastroVision, random folds contained 78–101 detected cross‐subset overlaps, whereas grouped folds contained none. Grouped EfficientNet‐B0 achieved 0.8195 ± 0.0043 accuracy and 0.5792 ± 0.0148 macro‐F1; ECE decreased from 0.1114 ± 0.0148 to 0.0323 ± 0.0048 after scaling. Similarity grouping removed detected resemblance without reducing accuracy, while calibration and uncertainty analyses exposed properties hidden by discrimination metrics. GastroVision supports methodological portability rather than direct external validation; pHash cannot establish patient or procedure independence, for which identifiers remain preferable.

Authors

Institutions

Publication Details

Journal
International Journal of Imaging Systems and Technology
Published
2026-10-06
DOI
https://doi.org/10.1002/ima.70452
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Similarity‐Aware Evaluation, Probability Calibration, and Independent Replication for Reliable Gastrointestinal Image Classification

Jinlian Zha, Yingwu Xu, Housheng Liu, Fang Cheng
International Journal of Imaging Systems and Technology
Advanced Neural Network Applications
article

Similarity‐Aware Evaluation, Probability Calibration, and Independent Replication for Reliable Gastrointestinal Image Classification

Jinlian Zha, Yingwu Xu, Housheng Liu, Fang Cheng
article en

Abstract

ABSTRACT High accuracy alone does not establish reliable medical‐image classification because visually similar images may cross partitions and confidence may be miscalibrated. We evaluated a similarity‐aware framework on a four‐class gastrointestinal benchmark (4000 images) and independently replicated the methodology on GastroVision (8000 images; 27 classes). A 64‐bit DCT perceptual hash defined groups at Hamming distance ≤ 4. Five‐fold random and group‐constrained protocols were matched by model and seed, with separate validation, calibration, and test subsets. We compared modern backbones, temperature scaling, a 2 × 2 Mixup/label‐smoothing ablation, and deterministic maximum‐softmax probability. On the primary benchmark, grouped EfficientNet‐B0 achieved accuracy and macro‐F1 of 0.9765 ± 0.0022. Grouping changed accuracy by +0.0025 relative to random splitting (Wilcoxon p = 0.625). Temperature scaling reduced expected calibration error from 0.1292 ± 0.0216 to 0.0123 ± 0.0030 without changing predictions, and deterministic MSP achieved error‐detection AUROC of 0.9288 ± 0.0181. On GastroVision, random folds contained 78–101 detected cross‐subset overlaps, whereas grouped folds contained none. Grouped EfficientNet‐B0 achieved 0.8195 ± 0.0043 accuracy and 0.5792 ± 0.0148 macro‐F1; ECE decreased from 0.1114 ± 0.0148 to 0.0323 ± 0.0048 after scaling. Similarity grouping removed detected resemblance without reducing accuracy, while calibration and uncertainty analyses exposed properties hidden by discrimination metrics. GastroVision supports methodological portability rather than direct external validation; pHash cannot establish patient or procedure independence, for which identifiers remain preferable.

International Journal of Imaging Systems and TechnologyVol. 36(6)
Anqing Normal University (CN)
Openalex Percentile: Top 15%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.