An uncertainty-aware ultrasound triage model for cervical lymphadenopathy: a retrospective multicenter development and external validation study

Binary artificial-intelligence classifiers do not indicate when an ultrasound case should be withheld from automatic routing. We evaluated a development-locked uncertainty-aware selective triage rule in a retrospective multicenter cohort. This retrospective multicenter study included 518 patients, comprising 206 patients in the development cohort, 88 in the internal-validation cohort, and 112 in each of two external-validation cohorts. One index lymph node was analyzed per patient. A fusion model integrated clinical, multimodal-ultrasound, and structured report-derived information. Preprocessing, feature selection, model fitting, selection of the conventional binary threshold ( p = 0.537) and the selective-routing probability boundaries ( p = 0.320 and p = 0.660), construction of composite uncertainty from development-normalized predictive entropy and bootstrap-based probability dispersion, and selection of the uncertainty threshold ( U = 0.700) were performed exclusively in the development cohort. The locked model and routing rule were subsequently applied without retuning to the internal- and external-validation cohorts. Evaluation included discrimination, calibration, paired comparator analyses, development-independent selective-triage performance, risk–coverage behavior, error detection, failure analysis, and an eight-reader study analyzed using crossed reader-by-case bootstrap resampling. The fusion-model area under the receiver operating characteristic curve (AUROC) was 0.839 in internal validation, 0.850 in external validation cohort 1, and 0.842 in external validation cohort 2. At the locked conventional binary threshold of p = 0.537, sensitivity/specificity were 0.595/0.902 in internal validation, 0.657/0.911 in external validation cohort 1, and 0.817/0.808 in external validation cohort 2. Development-independent coverage and selective accuracy were 0.364 and 0.906 in internal validation, 0.384 and 0.953 in external validation cohort 1, and 0.348 and 0.897 in external validation cohort 2, respectively. In the pooled external-validation population, 38 of 224 patients entered the lower-risk automatic pathway, 44 entered the higher-risk automatic pathway, and 142 were deferred for senior review. Automatic coverage was 0.366 (95% confidence interval [CI], 0.306–0.431), selective accuracy was 0.927 (95% CI, 0.849–0.966), and the accepted false-negative/false-positive counts were 2/4. Two of the 38 patients assigned to the lower-risk automatic pathway were malignant, corresponding to a negative predictive value (NPV) of 0.947 (95% CI, 0.827–0.985) and a residual malignancy risk of 0.053 (95% CI, 0.015–0.173). The higher-risk pathway had a positive predictive value (PPV) of 0.909 (95% CI, 0.788–0.964). The pooled external risk–coverage partial area under the risk-coverage curve (AURC) was 0.037. In a separate per-case error-ranking analysis, composite uncertainty detected binary-model errors with an AUROC of 0.581 and an area under the precision-recall curve (AUPRC) of 0.247. Junior-reader accuracy increased by 0.055 (95% CI, 0.022–0.089; p = 0.002), whereas the changes among middle-level and senior readers were not statistically significant. The locked strategy selectively routed approximately 37% of external cases but retained substantial deferral and accepted false-negative and false-positive errors. It should be interpreted as a retrospective, research-stage selective-routing framework rather than an established clinical safety or workflow tool.

Authors

Institutions

Publication Details

Journal
BMC Medical Imaging
Published
2026-09-12
DOI
https://doi.org/10.1186/s12880-026-02780-8
Primary Topic
Lymphadenopathy Diagnosis and Analysis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

An uncertainty-aware ultrasound triage model for cervical lymphadenopathy: a retrospective multicenter development and external validation study

Xiaochen Liu, Jing Ning, Hang Ling, Chenshan Dong et al.
BMC Medical Imaging
Lymphadenopathy Diagnosis and Analysis
article

An uncertainty-aware ultrasound triage model for cervical lymphadenopathy: a retrospective multicenter development and external validation study

Xiaochen Liu, Jing Ning, Hang Ling, Chenshan Dong, Ziwei Zhang, Cailing Lin
article en

Abstract

Binary artificial-intelligence classifiers do not indicate when an ultrasound case should be withheld from automatic routing. We evaluated a development-locked uncertainty-aware selective triage rule in a retrospective multicenter cohort. This retrospective multicenter study included 518 patients, comprising 206 patients in the development cohort, 88 in the internal-validation cohort, and 112 in each of two external-validation cohorts. One index lymph node was analyzed per patient. A fusion model integrated clinical, multimodal-ultrasound, and structured report-derived information. Preprocessing, feature selection, model fitting, selection of the conventional binary threshold ( p = 0.537) and the selective-routing probability boundaries ( p = 0.320 and p = 0.660), construction of composite uncertainty from development-normalized predictive entropy and bootstrap-based probability dispersion, and selection of the uncertainty threshold ( U = 0.700) were performed exclusively in the development cohort. The locked model and routing rule were subsequently applied without retuning to the internal- and external-validation cohorts. Evaluation included discrimination, calibration, paired comparator analyses, development-independent selective-triage performance, risk–coverage behavior, error detection, failure analysis, and an eight-reader study analyzed using crossed reader-by-case bootstrap resampling. The fusion-model area under the receiver operating characteristic curve (AUROC) was 0.839 in internal validation, 0.850 in external validation cohort 1, and 0.842 in external validation cohort 2. At the locked conventional binary threshold of p = 0.537, sensitivity/specificity were 0.595/0.902 in internal validation, 0.657/0.911 in external validation cohort 1, and 0.817/0.808 in external validation cohort 2. Development-independent coverage and selective accuracy were 0.364 and 0.906 in internal validation, 0.384 and 0.953 in external validation cohort 1, and 0.348 and 0.897 in external validation cohort 2, respectively. In the pooled external-validation population, 38 of 224 patients entered the lower-risk automatic pathway, 44 entered the higher-risk automatic pathway, and 142 were deferred for senior review. Automatic coverage was 0.366 (95% confidence interval [CI], 0.306–0.431), selective accuracy was 0.927 (95% CI, 0.849–0.966), and the accepted false-negative/false-positive counts were 2/4. Two of the 38 patients assigned to the lower-risk automatic pathway were malignant, corresponding to a negative predictive value (NPV) of 0.947 (95% CI, 0.827–0.985) and a residual malignancy risk of 0.053 (95% CI, 0.015–0.173). The higher-risk pathway had a positive predictive value (PPV) of 0.909 (95% CI, 0.788–0.964). The pooled external risk–coverage partial area under the risk-coverage curve (AURC) was 0.037. In a separate per-case error-ranking analysis, composite uncertainty detected binary-model errors with an AUROC of 0.581 and an area under the precision-recall curve (AUPRC) of 0.247. Junior-reader accuracy increased by 0.055 (95% CI, 0.022–0.089; p = 0.002), whereas the changes among middle-level and senior readers were not statistically significant. The locked strategy selectively routed approximately 37% of external cases but retained substantial deferral and accepted false-negative and false-positive errors. It should be interpreted as a retrospective, research-stage selective-routing framework rather than an established clinical safety or workflow tool.

BMC Medical Imaging
Fujian Medical University (CN), Fujian Provincial Hospital (CN)
Openalex Percentile: Top 8%
Lymphadenopathy Diagnosis and Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.