Clinically-guided self-refining large language model for automated diagnostic confirmation of suspected acute stroke in the emergency department: a retrospective development and internal validation study

Code stroke activations are widely implemented to accelerate diagnostic evaluation of suspected acute stroke in the emergency department. While effective in expediting care, false positive activations disrupt workflows, consume resources, and delay evaluation for other patients. We evaluated whether adaptations of a large language model could diagnostically confirm suspected acute stroke using information from initial emergency department documentation and compared results to BioClinical BERT. We analyzed initial ED clinical notes from Seoul National University Hospital (January 2018 to December 2022). Notes contained code-mixed (Korean and English), semi-structured and unstructured free text describing presenting illness, past medical history, and neurological deficits of patients presenting with acute neurological symptoms within 24 h of onset. LLaMA-3.1-70B-Instruct-GPTQ-INT4 was used to translate each text to English and augment negative labels in training folds. LLaMA-3.1-8B-Instruct was fine-tuned using (1) quantized low-rank adaptation and (2) a self-refinement instruction-tuning strategy applied with clinical rule of thumb augmentation. Performance was evaluated on the validation folds for AUROC, AUPRC, F1, accuracy, precision, recall, specificity, and Brier Score. Based on per-seed aggregation followed by average across seeds, BioClinical BERT yielded AUROC of 0.7230 ± 0.0200 (95% CI, 0.7011–0.7449) and Brier Score of 0.2081 ± 0.0049. Self-refinement, instruction-tuning of quantized low-rank adaptation with clinical rule of thumb augmentation yielded AUROC of 0.7882 ± 0.0102 (95% CI, 0.7770–0.7994) and Brier Score of 0.1577 ± 0.0020. A paired bootstrap comparison of AUROC with per-seed resampling (2000 iterations) and Fisher’s method for p-value aggregation demonstrated a statistically significant improvement ( p < 0.001). Self-refinement instruction-tuning combined with quantized low-rank adaptation and clinical rule of thumb guidance outperformed BioClinical BERT and offers a scalable approach for automated diagnostic confirmation of suspected acute stroke using initial emergency department documentation. External validation is still required prior to deployment.

Authors

Institutions

Publication Details

Journal
BMC Emergency Medicine
Published
2026-09-21
DOI
https://doi.org/10.1186/s12873-026-01788-1
Primary Topic
Acute Ischemic Stroke Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Clinically-guided self-refining large language model for automated diagnostic confirmation of suspected acute stroke in the emergency department: a retrospective development and internal validation study

Gil Joon Suh, Han‐Yeong Jeong, Jinwook Choi, Ji Han Heo
BMC Emergency Medicine
Acute Ischemic Stroke Management
article

Clinically-guided self-refining large language model for automated diagnostic confirmation of suspected acute stroke in the emergency department: a retrospective development and internal validation study

Gil Joon Suh, Han‐Yeong Jeong, Jinwook Choi, Ji Han Heo
article en

Abstract

Code stroke activations are widely implemented to accelerate diagnostic evaluation of suspected acute stroke in the emergency department. While effective in expediting care, false positive activations disrupt workflows, consume resources, and delay evaluation for other patients. We evaluated whether adaptations of a large language model could diagnostically confirm suspected acute stroke using information from initial emergency department documentation and compared results to BioClinical BERT. We analyzed initial ED clinical notes from Seoul National University Hospital (January 2018 to December 2022). Notes contained code-mixed (Korean and English), semi-structured and unstructured free text describing presenting illness, past medical history, and neurological deficits of patients presenting with acute neurological symptoms within 24 h of onset. LLaMA-3.1-70B-Instruct-GPTQ-INT4 was used to translate each text to English and augment negative labels in training folds. LLaMA-3.1-8B-Instruct was fine-tuned using (1) quantized low-rank adaptation and (2) a self-refinement instruction-tuning strategy applied with clinical rule of thumb augmentation. Performance was evaluated on the validation folds for AUROC, AUPRC, F1, accuracy, precision, recall, specificity, and Brier Score. Based on per-seed aggregation followed by average across seeds, BioClinical BERT yielded AUROC of 0.7230 ± 0.0200 (95% CI, 0.7011–0.7449) and Brier Score of 0.2081 ± 0.0049. Self-refinement, instruction-tuning of quantized low-rank adaptation with clinical rule of thumb augmentation yielded AUROC of 0.7882 ± 0.0102 (95% CI, 0.7770–0.7994) and Brier Score of 0.1577 ± 0.0020. A paired bootstrap comparison of AUROC with per-seed resampling (2000 iterations) and Fisher’s method for p-value aggregation demonstrated a statistically significant improvement ( p < 0.001). Self-refinement instruction-tuning combined with quantized low-rank adaptation and clinical rule of thumb guidance outperformed BioClinical BERT and offers a scalable approach for automated diagnostic confirmation of suspected acute stroke using initial emergency department documentation. External validation is still required prior to deployment.

BMC Emergency Medicine
Seoul National University (KR), Eulji University (KR), Seoul National University Hospital (KR), SNUH SMG-SNU Boramae Medical Center (KR), National Medical Center (KR)
Quality Education
Openalex Percentile: Top 11%
Acute Ischemic Stroke Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.