Low-burden identification of severe early deterioration among mid-acuity emergency department attendances initially triaged as category 3 using a retrieval-augmented large language model: a retrospective outcome-defined case-control study
Abstract Background Mid-acuity emergency department (ED) attendances are clinically heterogeneous; a small subset deteriorate early, but broad escalation would create substantial review burden. We evaluated whether retrieval augmentation improved low-burden risk enrichment versus a prompt-only large language model (LLM). Methods We conducted a retrospective outcome-defined case-control study nested within all 2025 nurse-triaged Category-3 attendances at a single Hong Kong ED. Cases had death within 24 h of ED registration or direct ICU/PICU admission; 472 endpoint-negative controls were selected using calendar-month–guided random sampling after prespecified exclusions and fixed before model inference, yielding a final case-to-control ratio of 1:4. Both DeepSeek-V3 configurations used identical triage-time inputs; MECR-RAG additionally retrieved local guideline sections and similar 2024 cases. Primary analyses compared linearly interpolated sensitivity at a 10% endpoint-negative alert burden and partial area under the receiver operating characteristic curve (pAUC, 0–0.10). Predicted Category ≤ 2 defined a separate fixed-threshold operational alert. NEWS2 was included as a post hoc exploratory clinical comparator when calculable. Results The cohort comprised 590 attendances (118 cases; 472 controls). Interpolated sensitivity at a 10% endpoint-negative alert burden was 68.1% (95% CI 56.5–76.3) for MECR-RAG and 27.8% (23.9–32.5) for baseline; pAUC was 0.0408 (0.0321–0.0508) and 0.0152 (0.0124–0.0186), respectively. At Category ≤ 2, MECR-RAG alerted 80/118 cases (67.8%) and 43/472 controls (9.1%), versus 104/118 (88.1%) and 160/472 (33.9%) for baseline, yielding a positive likelihood ratio of 7.44 (95% CI 5.45–10.16) versus 2.60 (2.26–3.00). A post hoc exploratory complete-case analysis included 280/590 attendances (47.5%); interpolated sensitivity at a 10% endpoint-negative alert burden was 58.3% for NEWS2 and 54.2% for MECR-RAG (difference − 4.0% points, 95% CI − 21.8 to 13.7). Conclusions Retrieval augmentation improved low-burden discrimination relative to a prompt-only LLM. It did not demonstrate superiority over NEWS2 in the post hoc exploratory complete-case comparison, and its incremental clinical value remains unproven. These results concern a narrow severe-outcome endpoint and do not establish triage correctness, clinical utility, or deployment readiness.
Authors
- Hang Sheung Wong (ORCID: https://orcid.org/0009-0003-4512-6547)
- TL Wong (ORCID: https://orcid.org/0009-0003-0042-4831)
- Ching Yan Li
Institutions
- Princess Margaret Hospital (HK)
Publication Details
- Journal
- BMC Emergency Medicine
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1186/s12873-026-01791-6
- Primary Topic
- Emergency and Acute Care Studies
- Type
- article
- Field-Weighted Citation Impact
- 0.00