Large language models fail to reliably predict emergent catheterization laboratory activation from prehospital electrocardiograms

Background Rapid and accurate electrocardiogram (ECG) interpretation is essential for timely identification of ST-elevation myocardial infarction (STEMI) and activation of reperfusion pathways in emergency care. Objectives To evaluate the diagnostic performance of multimodal LLMs in identifying prehospital ECGs warranting emergent catheterization laboratory activation. Methods We performed a retrospective analysis of 615 ECGs from 270 emergency medical service patient encounters (EMS) with concern for acute myocardial infarction. The reference standard was cardiology activation of the STEMI pathway for emergent angiography. LLM-based image interpretation (three models) and ECG machine algorithm interpretations were compared. Sensitivity, specificity, positive predictive value, negative predictive value, and overall accuracy were calculated. Results Gemini demonstrated the highest sensitivity (95.3%; 95% CI 91.7–97.3) but extremely poor specificity (9.4%), indicating a high false-positive rate. ChatGPT and Claude showed moderate sensitivity (68.1% and 67.2%) with limited specificity (42.3% and 46.5%). The ECG machine algorithm demonstrated more balanced performance, with sensitivity of 67.7% (95% CI 61.4–73.4) and higher specificity (64.2%) than all LLMs. Conclusions Multimodal LLM interpretation of prehospital ECGs demonstrated clinically unreliable performance for identifying ECGs warranting emergent cardiac catheterization laboratory activation when benchmarked against real-world cardiology activation decisions. Although some models achieved high sensitivity, poor specificity resulted in excessive false-positive activation recommendations. These findings suggest that general-purpose LLMs are not appropriate for ECG-based catheterization laboratory activation decisions in time-sensitive cardiopulmonary care workflows.

Authors

Institutions

Publication Details

Journal
Heart & Lung
Published
2026-09-16
DOI
https://doi.org/10.1016/j.hrtlng.2026.102956
Primary Topic
Cardiac Arrest and Resuscitation
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Large language models fail to reliably predict emergent catheterization laboratory activation from prehospital electrocardiograms

Urška Cvek, B. Watkins, Emile Legendre, Stewart Greathouse et al.
Heart & Lung
Cardiac Arrest and Resuscitation
article

Large language models fail to reliably predict emergent catheterization laboratory activation from prehospital electrocardiograms

Urška Cvek, B. Watkins, Emile Legendre, Stewart Greathouse, Dillon Jones, David Janese, Colton Toups
article en

Abstract

Background Rapid and accurate electrocardiogram (ECG) interpretation is essential for timely identification of ST-elevation myocardial infarction (STEMI) and activation of reperfusion pathways in emergency care. Objectives To evaluate the diagnostic performance of multimodal LLMs in identifying prehospital ECGs warranting emergent catheterization laboratory activation. Methods We performed a retrospective analysis of 615 ECGs from 270 emergency medical service patient encounters (EMS) with concern for acute myocardial infarction. The reference standard was cardiology activation of the STEMI pathway for emergent angiography. LLM-based image interpretation (three models) and ECG machine algorithm interpretations were compared. Sensitivity, specificity, positive predictive value, negative predictive value, and overall accuracy were calculated. Results Gemini demonstrated the highest sensitivity (95.3%; 95% CI 91.7–97.3) but extremely poor specificity (9.4%), indicating a high false-positive rate. ChatGPT and Claude showed moderate sensitivity (68.1% and 67.2%) with limited specificity (42.3% and 46.5%). The ECG machine algorithm demonstrated more balanced performance, with sensitivity of 67.7% (95% CI 61.4–73.4) and higher specificity (64.2%) than all LLMs. Conclusions Multimodal LLM interpretation of prehospital ECGs demonstrated clinically unreliable performance for identifying ECGs warranting emergent cardiac catheterization laboratory activation when benchmarked against real-world cardiology activation decisions. Although some models achieved high sensitivity, poor specificity resulted in excessive false-positive activation recommendations. These findings suggest that general-purpose LLMs are not appropriate for ECG-based catheterization laboratory activation decisions in time-sensitive cardiopulmonary care workflows.

Heart & LungVol. 80
Louisiana State University in Shreveport (US), Louisiana State University Health Sciences Center Shreveport (US)
No poverty
Openalex Percentile: Top 7%
Cardiac Arrest and Resuscitation
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.