Prediction of postoperative nausea and vomiting within 24 hours after total knee arthroplasty using interpretable machine learning

This study aimed to develop and evaluate an interpretable machine learning model for predicting postoperative nausea and vomiting (PONV) within 24 hours after total knee arthroplasty (TKA), and to compare its performance with the simplified Apfel score. We conducted a retrospective cohort study using routinely collected clinical data. Predictors were structured into three hierarchical feature sets: baseline demographic andcomorbidity variables (Step 1), PONV-related susceptibility factors (Step 2), and perioperative management variables (Step 3). Multiple machine learning models were developed and evaluated using cross-validation and a held-out internal test set. Model performance was assessed in terms of discrimination (area under the receiver operating characteristic curve [AUC] and area under the precision–recall curve [PR-AUC]), calibration (calibration plots and Brier score), and clinical utility using decision curve analysis. The simplified Apfel score was evaluated in the same test dataset for comparison. A temporal validation sensitivity analysis was additionally performed by fitting the final model in the earlier 80% of patients and evaluating it in the most recent 20%. A total of 1,188 patients were included, of whom 228 (19.2%) developed PONV. The random forest model achieved the best performance, with a test-set AUC of 0.711 and PR-AUC of 0.392. The simplified Apfel score showed moderate discrimination (AUC = 0.657). Calibration analysis demonstrated good agreement between predicted and observed probabilities for the machine learning model (Brier score = 0.144), whereas the Apfel score showed poor calibration with systematic underestimation of risk (Brier score = 0.432), reflecting its miscalibration relative to the observed event rate in this cohort. Decision curve analysis indicated that the machine learning model provided greater net benefit across clinically relevant threshold probabilities of approximately 10%–35%, while the Apfel score showed limited clinical utility. In the temporal validation cohort, the model achieved an AUC of 0.791, a PR-AUC of 0.451, and a Brier score of 0.124. The calibration intercept of 0.876 and slope of 1.970 indicated imperfect calibration of absolute risk estimates. An interpretable random forest model based on routinely available clinical data showed modest discrimination and more favorable calibration than the simplified Apfel score in the held-out internal test set. The temporal validation sensitivity analysis suggested thatdiscrimination was maintained in later-period patients, although absolute risk calibration remained imperfect. Further external multicenter and prospective validation, followed by evaluation of workflow integration and clinical impact, is required before routine clinical use can be considered.

Authors

Institutions

Publication Details

Journal
BMC Medical Informatics and Decision Making
Published
2026-09-11
DOI
https://doi.org/10.1186/s12911-026-03841-2
Primary Topic
Nausea and vomiting management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Prediction of postoperative nausea and vomiting within 24 hours after total knee arthroplasty using interpretable machine learning

Jae Bum Kwon, Inyoung Jung, Won Kee Choi, Sang Gyu Kwak et al.
BMC Medical Informatics and Decision Making
Nausea and vomiting management
article

Prediction of postoperative nausea and vomiting within 24 hours after total knee arthroplasty using interpretable machine learning

Jae Bum Kwon, Inyoung Jung, Won Kee Choi, Sang Gyu Kwak, Hyunsoo Kang
article en

Abstract

This study aimed to develop and evaluate an interpretable machine learning model for predicting postoperative nausea and vomiting (PONV) within 24 hours after total knee arthroplasty (TKA), and to compare its performance with the simplified Apfel score. We conducted a retrospective cohort study using routinely collected clinical data. Predictors were structured into three hierarchical feature sets: baseline demographic andcomorbidity variables (Step 1), PONV-related susceptibility factors (Step 2), and perioperative management variables (Step 3). Multiple machine learning models were developed and evaluated using cross-validation and a held-out internal test set. Model performance was assessed in terms of discrimination (area under the receiver operating characteristic curve [AUC] and area under the precision–recall curve [PR-AUC]), calibration (calibration plots and Brier score), and clinical utility using decision curve analysis. The simplified Apfel score was evaluated in the same test dataset for comparison. A temporal validation sensitivity analysis was additionally performed by fitting the final model in the earlier 80% of patients and evaluating it in the most recent 20%. A total of 1,188 patients were included, of whom 228 (19.2%) developed PONV. The random forest model achieved the best performance, with a test-set AUC of 0.711 and PR-AUC of 0.392. The simplified Apfel score showed moderate discrimination (AUC = 0.657). Calibration analysis demonstrated good agreement between predicted and observed probabilities for the machine learning model (Brier score = 0.144), whereas the Apfel score showed poor calibration with systematic underestimation of risk (Brier score = 0.432), reflecting its miscalibration relative to the observed event rate in this cohort. Decision curve analysis indicated that the machine learning model provided greater net benefit across clinically relevant threshold probabilities of approximately 10%–35%, while the Apfel score showed limited clinical utility. In the temporal validation cohort, the model achieved an AUC of 0.791, a PR-AUC of 0.451, and a Brier score of 0.124. The calibration intercept of 0.876 and slope of 1.970 indicated imperfect calibration of absolute risk estimates. An interpretable random forest model based on routinely available clinical data showed modest discrimination and more favorable calibration than the simplified Apfel score in the held-out internal test set. The temporal validation sensitivity analysis suggested thatdiscrimination was maintained in later-period patients, although absolute risk calibration remained imperfect. Further external multicenter and prospective validation, followed by evaluation of workflow integration and clinical impact, is required before routine clinical use can be considered.

BMC Medical Informatics and Decision Making
Daegu Catholic University (KR), Daegu Catholic University Medical Center (KR)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Nausea and vomiting management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.