Evaluating a Large Language Model for Emergency Triage Using Synthetic Clinical Dialogues: Toward a Multimodal Clinical Safety Framework

Background/Objectives: Prioritizing patient safety through accurate emergency triage is critical during mass-gathering events such as Hajj and Umrah, where high patient volumes can significantly increase the risk of clinical errors. Although the Emergency Severity Index (ESI) provides a standardized triage protocol, it remains dependent on human judgment, while traditional artificial intelligence (AI) models often struggle to capture the nuanced clinical context embedded in patient narratives. To address this gap, this study evaluates GPT-5.2 using synthetically generated dialogues as a step toward a conceptual clinical decision-support framework that bridges structured electronic health records (EHRs) with triage interactions through clinical dialogues. By combining narrative inputs with structured physiological indicators, it explores the potential to support more informed clinical decision-making under controlled simulation conditions. Methods: To evaluate this contribution, we systematically assessed the role of clinical dialogues in triage performance using a balanced and stratified dataset of 1000 simulated Hajj-context scenarios. Results: The results demonstrate that incorporating clinical dialogues can significantly improve triage performance, with a 79.7% relative reduction in under-triage across all ESI levels. A more noticeable improvement was observed in high-acuity cases, with a 91.9% relative reduction in under-triage. In addition, the model achieved a high-acuity recall of 94.5%, highlighting its ability to identify critical “red-flag” symptoms that are often not captured in structured data. These findings suggest that narrative context may contribute to more accurate triage assessment and improved identification of high-acuity cases. Conclusions: Consequently, when evaluated within a controlled decision-support setting, this study suggests that large language model (LLM)-based systems may improve triage reliability and may contribute to safer triage decision-making in high-demand clinical environments.

Authors

Institutions

Publication Details

Journal
Diagnostics
Published
2026-10-09
DOI
https://doi.org/10.3390/diagnostics16203268
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Evaluating a Large Language Model for Emergency Triage Using Synthetic Clinical Dialogues: Toward a Multimodal Clinical Safety Framework

Sultanah M. Alshammari, Reem Alotaibi, Fatimah Albalawi, Taha Masri
Diagnostics
Artificial Intelligence in Healthcare and Education
article

Evaluating a Large Language Model for Emergency Triage Using Synthetic Clinical Dialogues: Toward a Multimodal Clinical Safety Framework

Sultanah M. Alshammari, Reem Alotaibi, Fatimah Albalawi, Taha Masri
article en

Abstract

Background/Objectives: Prioritizing patient safety through accurate emergency triage is critical during mass-gathering events such as Hajj and Umrah, where high patient volumes can significantly increase the risk of clinical errors. Although the Emergency Severity Index (ESI) provides a standardized triage protocol, it remains dependent on human judgment, while traditional artificial intelligence (AI) models often struggle to capture the nuanced clinical context embedded in patient narratives. To address this gap, this study evaluates GPT-5.2 using synthetically generated dialogues as a step toward a conceptual clinical decision-support framework that bridges structured electronic health records (EHRs) with triage interactions through clinical dialogues. By combining narrative inputs with structured physiological indicators, it explores the potential to support more informed clinical decision-making under controlled simulation conditions. Methods: To evaluate this contribution, we systematically assessed the role of clinical dialogues in triage performance using a balanced and stratified dataset of 1000 simulated Hajj-context scenarios. Results: The results demonstrate that incorporating clinical dialogues can significantly improve triage performance, with a 79.7% relative reduction in under-triage across all ESI levels. A more noticeable improvement was observed in high-acuity cases, with a 91.9% relative reduction in under-triage. In addition, the model achieved a high-acuity recall of 94.5%, highlighting its ability to identify critical “red-flag” symptoms that are often not captured in structured data. These findings suggest that narrative context may contribute to more accurate triage assessment and improved identification of high-acuity cases. Conclusions: Consequently, when evaluated within a controlled decision-support setting, this study suggests that large language model (LLM)-based systems may improve triage reliability and may contribute to safer triage decision-making in high-demand clinical environments.

DiagnosticsVol. 16(20)
Vrije Universiteit Brussel (BE), King Abdulaziz University (SA), King Abdulaziz Hospital (SA)
Openalex Percentile: Top 19%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.