Machine Learning to Prioritize High-Severity Patient Safety Events for Institutional Investigation: Algorithm Development and Validation Study
Background Patient safety events (PSEs) are preventable incidents that cause, or have the potential to cause, harm to patients during their medical journey. Although incident reporting systems capture large volumes of such events, only a small proportion undergo comprehensive investigation due to the resource-intensive nature of the review process. Patient safety specialists are tasked with triaging PSEs to prioritize high-severity cases that warrant timely institutional investigation; however, the rapidly increasing volume of events has made manual triage increasingly impractical. Objective This study aimed to develop and validate ranking-based machine learning models for prioritizing high-severity PSEs and to compare their performance with conventional classification-based models and the reporter-assigned severity baseline. Methods A total of 101,239 PSE reports were retrospectively extracted from a large Canadian academic health system between January 2017 and March 2025. Each report included a free-text incident description and a corresponding severity label denoting the level of severity, assigned through an institutional review process. Both feature-based and transformer-based models were developed using ranking and classification frameworks. Model performance was evaluated using ranking metrics, including average precision and normalized discounted cumulative gain (NDCG) at multiple cutoff thresholds (k=10, 20, 50, 100, 200), as well as Precision@k and Recall@k. The best-performing model was further benchmarked against the reporter-assigned severity baseline. Results The ranking model, LLAMA3.1-RA, achieved the highest mean average precision of 0.94 (95% CI 0.90-0.98), representing a 22.1% relative improvement over the best classification model (LLAMA3.1-CL: mean 0.77, 95% CI 0.72-0.82; P<.001). LLAMA3.1-RA demonstrated superior performance across 10 of 16 evaluation metrics, with larger relative gains at broader evaluation cutoffs (eg, +12.9% in NDCG@200, +21.1% in Precision@200, and +20.3% in Recall@200). Compared with the reporter-assigned severity baseline, LLAMA3.1-RA achieved significantly higher performance across all 16 metrics, retrieving 83.0% of high-severity PSEs within the top 200 ranked events, compared with 41.0% retrieved under the reporter-assigned severity baseline. Conclusions LLAMA3.1-RA demonstrated superior capability in prioritizing high-severity PSEs compared with both classification approaches and the reporter-assigned severity baseline. Integrating this model into triage workflows may offer a scalable and efficient solution for identifying and prioritizing high-severity safety events.
Authors
- Lucas Brien Chartier (ORCID: https://orcid.org/0000-0001-9716-1684)
- Shehnaz Islam (ORCID: https://orcid.org/0000-0002-4032-1445)
- Laura Danielle Pozzobon (ORCID: https://orcid.org/0000-0002-1321-8872)
- Eldan Cohen (ORCID: https://orcid.org/0000-0001-5767-6683)
- Hongbo Chen (ORCID: https://orcid.org/0009-0005-5823-9406)
Publication Details
- Journal
- JMIR Medical Informatics
- Published
- 2026-09-29
- DOI
- https://doi.org/10.2196/101393
- Primary Topic
- Patient Safety and Medication Errors
- Type
- article
- Field-Weighted Citation Impact
- 0.00