A medically grounded LLM agent–based tool to detect patient safety events in medical records

Large language models (LLMs) have shown incredible promise in medicine. While LLMs may be particularly useful in areas requiring extensive review of clinical records, their use remains limited due to their tendency to hallucinate and fabricate information. Hallucination issues, as well as their consequences, are exacerbated in low-probability, high-stakes scenarios such as rare adverse safety events or medical errors. We present SAFE-AI (Structured and Automated Framework for Explainable AI), a novel method for clinical decision making that combines the strengths of clinical expert knowledge with LLMs in an ontology-driven model that minimizes hallucinations using strict rules. We test this method to identify medication errors in medical charts. We collected a sample of 18,402 lines of clinical information from 300 EMS clinical charts that were independently dually reviewed by two expert physicians for epinephrine adverse safety events (ASEs), with 96% inter-rater agreement. We tested SAFE-AI against these labels, achieving human-like performance in detecting epinephrine overdoses with 97.9% accuracy, and 91.6% accuracy in identifying delays in epinephrine administration, greatly outperforming baseline LLMs models. Notably, some disagreements between clinicians and the model were found to be justifiable differences in judgment rather than errors. SAFE-AI presents a novel approach for clinical AI applications that addresses two key limitations of current machine learning methods: 1) over-reliance on probabilistic pattern recognition instead of established medical knowledge, and 2) perpetuation of biases present in training data. This framework is easily adaptable to a range of clinical applications, paving the way for provable and trustworthy AI in medicine.

Authors

Institutions

Publication Details

Journal
medRxiv
Published
2025-12-18
DOI
https://doi.org/10.64898/2025.12.16.25342438
Citations
1
Primary Topic
Machine Learning in Healthcare
Type
article
Field-Weighted Citation Impact
2.48
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A medically grounded LLM agent–based tool to detect patient safety events in medical records

David Alexander Vernaza Trujillo, Nathan Bahr, K. Kim, Byeongyeon Cho et al.
1 citations
medRxiv
Machine Learning in Healthcare
2.48
article

A medically grounded LLM agent–based tool to detect patient safety events in medical records

David Alexander Vernaza Trujillo, Nathan Bahr, K. Kim, Byeongyeon Cho, Garth Meckler, Steven Bedrick, Dulin Wang, Xiaoqian Jiang, Matthew Hansen, Carl Eriksson, Tina Yi Jin Hsieh, Jeanne‐Marie Guise
article en
1 citations

Abstract

Large language models (LLMs) have shown incredible promise in medicine. While LLMs may be particularly useful in areas requiring extensive review of clinical records, their use remains limited due to their tendency to hallucinate and fabricate information. Hallucination issues, as well as their consequences, are exacerbated in low-probability, high-stakes scenarios such as rare adverse safety events or medical errors. We present SAFE-AI (Structured and Automated Framework for Explainable AI), a novel method for clinical decision making that combines the strengths of clinical expert knowledge with LLMs in an ontology-driven model that minimizes hallucinations using strict rules. We test this method to identify medication errors in medical charts. We collected a sample of 18,402 lines of clinical information from 300 EMS clinical charts that were independently dually reviewed by two expert physicians for epinephrine adverse safety events (ASEs), with 96% inter-rater agreement. We tested SAFE-AI against these labels, achieving human-like performance in detecting epinephrine overdoses with 97.9% accuracy, and 91.6% accuracy in identifying delays in epinephrine administration, greatly outperforming baseline LLMs models. Notably, some disagreements between clinicians and the model were found to be justifiable differences in judgment rather than errors. SAFE-AI presents a novel approach for clinical AI applications that addresses two key limitations of current machine learning methods: 1) over-reliance on probabilistic pattern recognition instead of established medical knowledge, and 2) perpetuation of biases present in training data. This framework is easily adaptable to a range of clinical applications, paving the way for provable and trustworthy AI in medicine.

medRxiv
Beth Israel Deaconess Medical Center (US), Harvard University (US), University of British Columbia (CA), Oregon Health & Science University (US), Hadassah Medical Center (IL), The University of Texas Health Science Center (US), The University of Texas Health Science Center at Houston (US)
Openalex Percentile: Top 8%
Machine Learning in Healthcare
2.48
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.