Device-Specific Conformal Reliability under Scanner-Level Acquisition Shift in Breast Ultrasound CAD

Conformal prediction provides finite-sample coverage control under exchangeability, but empirical reliability may change across scanner-defined deployment domains. We audit three conformal procedures and one non-conformal abstention baseline on leave-one-device-out (LODO) partitions of BUS-BRA using EfficientNet-B0, ResNet-50 and ConvNeXt-Tiny. All conformal thresholds are computed by direct finite-sample order-statistic indexing. The primary prediction and score unit is the image; source splits are patient-disjoint, target calibration/test partitions are patient-disjoint within each recalibration draw, and patient-cluster bootstrap plus patient-aggregated analyses assess within-patient dependence. At nominal α=0.05, EfficientNet-B0 Mondrian CP gives malignant set-miss rates of 0.004 for GE Logiq 7, 0.002 for GE Logiq 5 and 0.244 for Toshiba; the Toshiba patient-cluster 95% interval is 0.140-0.359. Patient aggregation preserves the contrast (0.194 for Toshiba versus 0.000 for both GE devices). In paired target recalibration, k=10 lowers Toshiba mean miss rate from 0.246 to 0.057 while increasing abstention to 0.603, whereas the same intervention raises the GE means to 0.060 and 0.067. Across the 500 Toshiba k=10 seed-draw combinations, the post-recalibration median miss rate is 0.021 (IQR 0.000-0.063), and 96.4% of draws improve relative to their paired pre-recalibration value. A finite-resolution analysis shows that α=0.02 is rank-unavailable for all tested few-shot conditions; under direct finite-sample order-statistic indexing, α=0.02 and α=0.05 coincide at k=5 and k=10 because both select the observed class maxima, but they separate in the larger k=20 condition. These results document scanner-level differences in empirical conformal reliability and motivate evaluating target recalibration when an independent target-domain reliability audit provides evidence of undercoverage, rather than establishing a prospective trigger rule.

Authors

Institutions

Publication Details

Journal
Medical Engineering & Physics
Published
2026-10-05
DOI
https://doi.org/10.1088/1873-4030/aeb029
Primary Topic
AI in cancer detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Device-Specific Conformal Reliability under Scanner-Level Acquisition Shift in Breast Ultrasound CAD

İsmail Hakkı Kinalioğlu
Medical Engineering & Physics
AI in cancer detection
article

Device-Specific Conformal Reliability under Scanner-Level Acquisition Shift in Breast Ultrasound CAD

İsmail Hakkı Kinalioğlu
article en

Abstract

Conformal prediction provides finite-sample coverage control under exchangeability, but empirical reliability may change across scanner-defined deployment domains. We audit three conformal procedures and one non-conformal abstention baseline on leave-one-device-out (LODO) partitions of BUS-BRA using EfficientNet-B0, ResNet-50 and ConvNeXt-Tiny. All conformal thresholds are computed by direct finite-sample order-statistic indexing. The primary prediction and score unit is the image; source splits are patient-disjoint, target calibration/test partitions are patient-disjoint within each recalibration draw, and patient-cluster bootstrap plus patient-aggregated analyses assess within-patient dependence. At nominal α=0.05, EfficientNet-B0 Mondrian CP gives malignant set-miss rates of 0.004 for GE Logiq 7, 0.002 for GE Logiq 5 and 0.244 for Toshiba; the Toshiba patient-cluster 95% interval is 0.140-0.359. Patient aggregation preserves the contrast (0.194 for Toshiba versus 0.000 for both GE devices). In paired target recalibration, k=10 lowers Toshiba mean miss rate from 0.246 to 0.057 while increasing abstention to 0.603, whereas the same intervention raises the GE means to 0.060 and 0.067. Across the 500 Toshiba k=10 seed-draw combinations, the post-recalibration median miss rate is 0.021 (IQR 0.000-0.063), and 96.4% of draws improve relative to their paired pre-recalibration value. A finite-resolution analysis shows that α=0.02 is rank-unavailable for all tested few-shot conditions; under direct finite-sample order-statistic indexing, α=0.02 and α=0.05 coincide at k=5 and k=10 because both select the observed class maxima, but they separate in the larger k=20 condition. These results document scanner-level differences in empirical conformal reliability and motivate evaluating target recalibration when an independent target-domain reliability audit provides evidence of undercoverage, rather than establishing a prospective trigger rule.

Medical Engineering & Physics
Selçuk University (TR)
Good health and well-being
Openalex Percentile: Top 12%
AI in cancer detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.