Machine Learning to Prioritize High-Severity Patient Safety Events for Institutional Investigation: Algorithm Development and Validation Study

Background Patient safety events (PSEs) are preventable incidents that cause, or have the potential to cause, harm to patients during their medical journey. Although incident reporting systems capture large volumes of such events, only a small proportion undergo comprehensive investigation due to the resource-intensive nature of the review process. Patient safety specialists are tasked with triaging PSEs to prioritize high-severity cases that warrant timely institutional investigation; however, the rapidly increasing volume of events has made manual triage increasingly impractical. Objective This study aimed to develop and validate ranking-based machine learning models for prioritizing high-severity PSEs and to compare their performance with conventional classification-based models and the reporter-assigned severity baseline. Methods A total of 101,239 PSE reports were retrospectively extracted from a large Canadian academic health system between January 2017 and March 2025. Each report included a free-text incident description and a corresponding severity label denoting the level of severity, assigned through an institutional review process. Both feature-based and transformer-based models were developed using ranking and classification frameworks. Model performance was evaluated using ranking metrics, including average precision and normalized discounted cumulative gain (NDCG) at multiple cutoff thresholds (k=10, 20, 50, 100, 200), as well as Precision@k and Recall@k. The best-performing model was further benchmarked against the reporter-assigned severity baseline. Results The ranking model, LLAMA3.1-RA, achieved the highest mean average precision of 0.94 (95% CI 0.90-0.98), representing a 22.1% relative improvement over the best classification model (LLAMA3.1-CL: mean 0.77, 95% CI 0.72-0.82; P<.001). LLAMA3.1-RA demonstrated superior performance across 10 of 16 evaluation metrics, with larger relative gains at broader evaluation cutoffs (eg, +12.9% in NDCG@200, +21.1% in Precision@200, and +20.3% in Recall@200). Compared with the reporter-assigned severity baseline, LLAMA3.1-RA achieved significantly higher performance across all 16 metrics, retrieving 83.0% of high-severity PSEs within the top 200 ranked events, compared with 41.0% retrieved under the reporter-assigned severity baseline. Conclusions LLAMA3.1-RA demonstrated superior capability in prioritizing high-severity PSEs compared with both classification approaches and the reporter-assigned severity baseline. Integrating this model into triage workflows may offer a scalable and efficient solution for identifying and prioritizing high-severity safety events.

Authors

Publication Details

Journal
JMIR Medical Informatics
Published
2026-09-29
DOI
https://doi.org/10.2196/101393
Primary Topic
Patient Safety and Medication Errors
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Machine Learning to Prioritize High-Severity Patient Safety Events for Institutional Investigation: Algorithm Development and Validation Study

Lucas Brien Chartier, Shehnaz Islam, Laura Danielle Pozzobon, Eldan Cohen et al.
JMIR Medical Informatics
Patient Safety and Medication Errors
article

Machine Learning to Prioritize High-Severity Patient Safety Events for Institutional Investigation: Algorithm Development and Validation Study

Lucas Brien Chartier, Shehnaz Islam, Laura Danielle Pozzobon, Eldan Cohen, Hongbo Chen
article en

Abstract

Background Patient safety events (PSEs) are preventable incidents that cause, or have the potential to cause, harm to patients during their medical journey. Although incident reporting systems capture large volumes of such events, only a small proportion undergo comprehensive investigation due to the resource-intensive nature of the review process. Patient safety specialists are tasked with triaging PSEs to prioritize high-severity cases that warrant timely institutional investigation; however, the rapidly increasing volume of events has made manual triage increasingly impractical. Objective This study aimed to develop and validate ranking-based machine learning models for prioritizing high-severity PSEs and to compare their performance with conventional classification-based models and the reporter-assigned severity baseline. Methods A total of 101,239 PSE reports were retrospectively extracted from a large Canadian academic health system between January 2017 and March 2025. Each report included a free-text incident description and a corresponding severity label denoting the level of severity, assigned through an institutional review process. Both feature-based and transformer-based models were developed using ranking and classification frameworks. Model performance was evaluated using ranking metrics, including average precision and normalized discounted cumulative gain (NDCG) at multiple cutoff thresholds (k=10, 20, 50, 100, 200), as well as Precision@k and Recall@k. The best-performing model was further benchmarked against the reporter-assigned severity baseline. Results The ranking model, LLAMA3.1-RA, achieved the highest mean average precision of 0.94 (95% CI 0.90-0.98), representing a 22.1% relative improvement over the best classification model (LLAMA3.1-CL: mean 0.77, 95% CI 0.72-0.82; P<.001). LLAMA3.1-RA demonstrated superior performance across 10 of 16 evaluation metrics, with larger relative gains at broader evaluation cutoffs (eg, +12.9% in NDCG@200, +21.1% in Precision@200, and +20.3% in Recall@200). Compared with the reporter-assigned severity baseline, LLAMA3.1-RA achieved significantly higher performance across all 16 metrics, retrieving 83.0% of high-severity PSEs within the top 200 ranked events, compared with 41.0% retrieved under the reporter-assigned severity baseline. Conclusions LLAMA3.1-RA demonstrated superior capability in prioritizing high-severity PSEs compared with both classification approaches and the reporter-assigned severity baseline. Integrating this model into triage workflows may offer a scalable and efficient solution for identifying and prioritizing high-severity safety events.

JMIR Medical InformaticsVol. 14
Openalex Percentile: Top 8%
Patient Safety and Medication Errors
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.