Low-burden identification of severe early deterioration among mid-acuity emergency department attendances initially triaged as category 3 using a retrieval-augmented large language model: a retrospective outcome-defined case-control study

Abstract Background Mid-acuity emergency department (ED) attendances are clinically heterogeneous; a small subset deteriorate early, but broad escalation would create substantial review burden. We evaluated whether retrieval augmentation improved low-burden risk enrichment versus a prompt-only large language model (LLM). Methods We conducted a retrospective outcome-defined case-control study nested within all 2025 nurse-triaged Category-3 attendances at a single Hong Kong ED. Cases had death within 24 h of ED registration or direct ICU/PICU admission; 472 endpoint-negative controls were selected using calendar-month–guided random sampling after prespecified exclusions and fixed before model inference, yielding a final case-to-control ratio of 1:4. Both DeepSeek-V3 configurations used identical triage-time inputs; MECR-RAG additionally retrieved local guideline sections and similar 2024 cases. Primary analyses compared linearly interpolated sensitivity at a 10% endpoint-negative alert burden and partial area under the receiver operating characteristic curve (pAUC, 0–0.10). Predicted Category ≤ 2 defined a separate fixed-threshold operational alert. NEWS2 was included as a post hoc exploratory clinical comparator when calculable. Results The cohort comprised 590 attendances (118 cases; 472 controls). Interpolated sensitivity at a 10% endpoint-negative alert burden was 68.1% (95% CI 56.5–76.3) for MECR-RAG and 27.8% (23.9–32.5) for baseline; pAUC was 0.0408 (0.0321–0.0508) and 0.0152 (0.0124–0.0186), respectively. At Category ≤ 2, MECR-RAG alerted 80/118 cases (67.8%) and 43/472 controls (9.1%), versus 104/118 (88.1%) and 160/472 (33.9%) for baseline, yielding a positive likelihood ratio of 7.44 (95% CI 5.45–10.16) versus 2.60 (2.26–3.00). A post hoc exploratory complete-case analysis included 280/590 attendances (47.5%); interpolated sensitivity at a 10% endpoint-negative alert burden was 58.3% for NEWS2 and 54.2% for MECR-RAG (difference − 4.0% points, 95% CI − 21.8 to 13.7). Conclusions Retrieval augmentation improved low-burden discrimination relative to a prompt-only LLM. It did not demonstrate superiority over NEWS2 in the post hoc exploratory complete-case comparison, and its incremental clinical value remains unproven. These results concern a narrow severe-outcome endpoint and do not establish triage correctness, clinical utility, or deployment readiness.

Authors

Institutions

Publication Details

Journal
BMC Emergency Medicine
Published
2026-09-24
DOI
https://doi.org/10.1186/s12873-026-01791-6
Primary Topic
Emergency and Acute Care Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Low-burden identification of severe early deterioration among mid-acuity emergency department attendances initially triaged as category 3 using a retrieval-augmented large language model: a retrospective outcome-defined case-control study

Hang Sheung Wong, TL Wong, Ching Yan Li
BMC Emergency Medicine
Emergency and Acute Care Studies
article

Low-burden identification of severe early deterioration among mid-acuity emergency department attendances initially triaged as category 3 using a retrieval-augmented large language model: a retrospective outcome-defined case-control study

Hang Sheung Wong, TL Wong, Ching Yan Li
article en

Abstract

Abstract Background Mid-acuity emergency department (ED) attendances are clinically heterogeneous; a small subset deteriorate early, but broad escalation would create substantial review burden. We evaluated whether retrieval augmentation improved low-burden risk enrichment versus a prompt-only large language model (LLM). Methods We conducted a retrospective outcome-defined case-control study nested within all 2025 nurse-triaged Category-3 attendances at a single Hong Kong ED. Cases had death within 24 h of ED registration or direct ICU/PICU admission; 472 endpoint-negative controls were selected using calendar-month–guided random sampling after prespecified exclusions and fixed before model inference, yielding a final case-to-control ratio of 1:4. Both DeepSeek-V3 configurations used identical triage-time inputs; MECR-RAG additionally retrieved local guideline sections and similar 2024 cases. Primary analyses compared linearly interpolated sensitivity at a 10% endpoint-negative alert burden and partial area under the receiver operating characteristic curve (pAUC, 0–0.10). Predicted Category ≤ 2 defined a separate fixed-threshold operational alert. NEWS2 was included as a post hoc exploratory clinical comparator when calculable. Results The cohort comprised 590 attendances (118 cases; 472 controls). Interpolated sensitivity at a 10% endpoint-negative alert burden was 68.1% (95% CI 56.5–76.3) for MECR-RAG and 27.8% (23.9–32.5) for baseline; pAUC was 0.0408 (0.0321–0.0508) and 0.0152 (0.0124–0.0186), respectively. At Category ≤ 2, MECR-RAG alerted 80/118 cases (67.8%) and 43/472 controls (9.1%), versus 104/118 (88.1%) and 160/472 (33.9%) for baseline, yielding a positive likelihood ratio of 7.44 (95% CI 5.45–10.16) versus 2.60 (2.26–3.00). A post hoc exploratory complete-case analysis included 280/590 attendances (47.5%); interpolated sensitivity at a 10% endpoint-negative alert burden was 58.3% for NEWS2 and 54.2% for MECR-RAG (difference − 4.0% points, 95% CI − 21.8 to 13.7). Conclusions Retrieval augmentation improved low-burden discrimination relative to a prompt-only LLM. It did not demonstrate superiority over NEWS2 in the post hoc exploratory complete-case comparison, and its incremental clinical value remains unproven. These results concern a narrow severe-outcome endpoint and do not establish triage correctness, clinical utility, or deployment readiness.

BMC Emergency Medicine
Princess Margaret Hospital (HK)
Quality Education, Reduced inequalities
Openalex Percentile: Top 8%
Emergency and Acute Care Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.