Operating-Point Collapse Under Normal-Class Source Confounding in Cross-Dataset Breast Ultrasound Classification

Background: Public breast ultrasound datasets are assembled around lesions. Of eight sources examined here, five contain no normal images, and under leave-one-dataset-out evaluation, the normal class is therefore almost perfectly confounded with the acquisition source: with one site held out, 356 of the 358 normal training images come from a single other source, supplied by 28 of that source’s 38 patients. Methods: We evaluated normal-versus-abnormal classification across eight public datasets with splits disjoint in the grouping identifier that each source supplies, two architectures and ten seeds per cell, treating the seed as the unit of analysis and comparing arms paired by seed. Seven interventions were compared: loss reweighting; balanced sampling; focal loss; class-balanced loss; and three variants of Mosaic, a patch substitution scheme that replaces annotated lesions with real non-lesion tissue. Results: Baseline models attained a mean area under the curve of 0.820 on held-out sites while assigning almost every image to the abnormal class (specificity 0.052 at threshold 0.5; 33 of 40 runs below 0.05). Validation area under the curve was 0.997, so epoch and threshold selection chose among near-ties, and identical reruns differed by 0.204 in specificity at the validation-selected threshold. A threshold set for 95% validation sensitivity delivered 74–92% on held-out sites. Subsampling the normal training images to 10% changed nothing, suggesting that volume is not the binding constraint and consistent with a role for source diversity. Only patch substitution improved the operating point, moving the mean specificity at threshold 0.5 from 0.052 to 0.235 and the expected calibration error from 0.348 to 0.234. Conclusions: In this setting, an operating point is not a property of the method, and results should be reported at a declared fixed threshold with per-seed variability.

Authors

Institutions

Publication Details

Journal
Information
Published
2026-09-28
DOI
https://doi.org/10.3390/info17100954
Primary Topic
Ultrasound Imaging and Elastography
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Operating-Point Collapse Under Normal-Class Source Confounding in Cross-Dataset Breast Ultrasound Classification

Beibit Abdikenov, Tomiris Zhaksylyk, Aruzhan Imasheva, Dauren Izdibay
Information
Ultrasound Imaging and Elastography
article

Operating-Point Collapse Under Normal-Class Source Confounding in Cross-Dataset Breast Ultrasound Classification

Beibit Abdikenov, Tomiris Zhaksylyk, Aruzhan Imasheva, Dauren Izdibay
article en

Abstract

Background: Public breast ultrasound datasets are assembled around lesions. Of eight sources examined here, five contain no normal images, and under leave-one-dataset-out evaluation, the normal class is therefore almost perfectly confounded with the acquisition source: with one site held out, 356 of the 358 normal training images come from a single other source, supplied by 28 of that source’s 38 patients. Methods: We evaluated normal-versus-abnormal classification across eight public datasets with splits disjoint in the grouping identifier that each source supplies, two architectures and ten seeds per cell, treating the seed as the unit of analysis and comparing arms paired by seed. Seven interventions were compared: loss reweighting; balanced sampling; focal loss; class-balanced loss; and three variants of Mosaic, a patch substitution scheme that replaces annotated lesions with real non-lesion tissue. Results: Baseline models attained a mean area under the curve of 0.820 on held-out sites while assigning almost every image to the abnormal class (specificity 0.052 at threshold 0.5; 33 of 40 runs below 0.05). Validation area under the curve was 0.997, so epoch and threshold selection chose among near-ties, and identical reruns differed by 0.204 in specificity at the validation-selected threshold. A threshold set for 95% validation sensitivity delivered 74–92% on held-out sites. Subsampling the normal training images to 10% changed nothing, suggesting that volume is not the binding constraint and consistent with a role for source diversity. Only patch substitution improved the operating point, moving the mean specificity at threshold 0.5 from 0.052 to 0.235 and the expected calibration error from 0.348 to 0.234. Conclusions: In this setting, an operating point is not a property of the method, and results should be reported at a declared fixed threshold with per-seed variability.

InformationVol. 17(10)
Astana IT University (KZ)
Openalex Percentile: Top 12%
Ultrasound Imaging and Elastography
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.