Prediction of postoperative nausea and vomiting within 24 hours after total knee arthroplasty using interpretable machine learning
This study aimed to develop and evaluate an interpretable machine learning model for predicting postoperative nausea and vomiting (PONV) within 24 hours after total knee arthroplasty (TKA), and to compare its performance with the simplified Apfel score. We conducted a retrospective cohort study using routinely collected clinical data. Predictors were structured into three hierarchical feature sets: baseline demographic andcomorbidity variables (Step 1), PONV-related susceptibility factors (Step 2), and perioperative management variables (Step 3). Multiple machine learning models were developed and evaluated using cross-validation and a held-out internal test set. Model performance was assessed in terms of discrimination (area under the receiver operating characteristic curve [AUC] and area under the precision–recall curve [PR-AUC]), calibration (calibration plots and Brier score), and clinical utility using decision curve analysis. The simplified Apfel score was evaluated in the same test dataset for comparison. A temporal validation sensitivity analysis was additionally performed by fitting the final model in the earlier 80% of patients and evaluating it in the most recent 20%. A total of 1,188 patients were included, of whom 228 (19.2%) developed PONV. The random forest model achieved the best performance, with a test-set AUC of 0.711 and PR-AUC of 0.392. The simplified Apfel score showed moderate discrimination (AUC = 0.657). Calibration analysis demonstrated good agreement between predicted and observed probabilities for the machine learning model (Brier score = 0.144), whereas the Apfel score showed poor calibration with systematic underestimation of risk (Brier score = 0.432), reflecting its miscalibration relative to the observed event rate in this cohort. Decision curve analysis indicated that the machine learning model provided greater net benefit across clinically relevant threshold probabilities of approximately 10%–35%, while the Apfel score showed limited clinical utility. In the temporal validation cohort, the model achieved an AUC of 0.791, a PR-AUC of 0.451, and a Brier score of 0.124. The calibration intercept of 0.876 and slope of 1.970 indicated imperfect calibration of absolute risk estimates. An interpretable random forest model based on routinely available clinical data showed modest discrimination and more favorable calibration than the simplified Apfel score in the held-out internal test set. The temporal validation sensitivity analysis suggested thatdiscrimination was maintained in later-period patients, although absolute risk calibration remained imperfect. Further external multicenter and prospective validation, followed by evaluation of workflow integration and clinical impact, is required before routine clinical use can be considered.
Authors
- Jae Bum Kwon
- Inyoung Jung (ORCID: https://orcid.org/0000-0002-6221-9177)
- Won Kee Choi
- Sang Gyu Kwak (ORCID: https://orcid.org/0000-0003-0398-5514)
- Hyunsoo Kang
Institutions
- Daegu Catholic University (KR)
- Daegu Catholic University Medical Center (KR)
Publication Details
- Journal
- BMC Medical Informatics and Decision Making
- Published
- 2026-09-11
- DOI
- https://doi.org/10.1186/s12911-026-03841-2
- Primary Topic
- Nausea and vomiting management
- Type
- article
- Field-Weighted Citation Impact
- 0.00