Repeated Internal Validation and Interpretation of Models Predicting Surgery Documented at a Study Institution Within 90 Days After Acute Knee Trauma: A Secondary Analysis

Background and Objectives: A previous study using this cohort evaluated threshold-dependent performance and clinical decision utility for predicting surgery after acute knee trauma. This secondary analysis assessed the stability of discrimination, probabilistic accuracy, calibration, model ranking, and predictor importance under repeated internal validation. Materials and Methods: We analyzed 905 adults with acute knee trauma, including 163 (18.0%) who had a qualifying surgical procedure documented at the study institution within 90 days of the index emergency department visit. Logistic regression, random forest, and gradient-boosted trees were evaluated using 50 repetitions of stratified five-fold cross-validation with identical folds across models. Each patient’s 50 out-of-fold predictions were averaged. Confidence intervals and paired model differences were estimated using 2000 patient-level stratified bootstrap resamples. Random-forest predictor importance was evaluated by validation-fold permutation. Results: Logistic regression and random forest had AUROCs of 0.762 and 0.757, AUPRCs of 0.430 and 0.443, and Brier scores of 0.125 and 0.123, respectively. Paired comparisons between these two models did not indicate statistically significant differences in any of the three metrics. Random forest had a higher AUPRC and lower Brier score than gradient-boosted trees. Calibration-in-the-large was close to 0 for all three models, whereas the calibration slope was compatible with 1 for logistic regression and random forest but was below 1 for gradient-boosted trees. Knee swelling, limited range of motion, and injury mechanism were the most influential random-forest predictors. Conclusions: Repeated internal validation did not identify a uniformly superior model. Logistic regression most often ranked first for AUROC, whereas random forest most often ranked first for AUPRC and Brier score; however, paired comparisons did not identify significant differences between these two models. External validation is required before clinical use.

Authors

Institutions

Publication Details

Journal
Medicina
Published
2026-09-22
DOI
https://doi.org/10.3390/medicina62101821
Primary Topic
Total Knee Arthroplasty Outcomes
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Repeated Internal Validation and Interpretation of Models Predicting Surgery Documented at a Study Institution Within 90 Days After Acute Knee Trauma: A Secondary Analysis

Joung Eun Lee, Won-Kee Choi, Sang Gyu Kwak, Young Woo Seo et al.
Medicina
Total Knee Arthroplasty Outcomes
article

Repeated Internal Validation and Interpretation of Models Predicting Surgery Documented at a Study Institution Within 90 Days After Acute Knee Trauma: A Secondary Analysis

Joung Eun Lee, Won-Kee Choi, Sang Gyu Kwak, Young Woo Seo, Seongho Hwang
article en

Abstract

Background and Objectives: A previous study using this cohort evaluated threshold-dependent performance and clinical decision utility for predicting surgery after acute knee trauma. This secondary analysis assessed the stability of discrimination, probabilistic accuracy, calibration, model ranking, and predictor importance under repeated internal validation. Materials and Methods: We analyzed 905 adults with acute knee trauma, including 163 (18.0%) who had a qualifying surgical procedure documented at the study institution within 90 days of the index emergency department visit. Logistic regression, random forest, and gradient-boosted trees were evaluated using 50 repetitions of stratified five-fold cross-validation with identical folds across models. Each patient’s 50 out-of-fold predictions were averaged. Confidence intervals and paired model differences were estimated using 2000 patient-level stratified bootstrap resamples. Random-forest predictor importance was evaluated by validation-fold permutation. Results: Logistic regression and random forest had AUROCs of 0.762 and 0.757, AUPRCs of 0.430 and 0.443, and Brier scores of 0.125 and 0.123, respectively. Paired comparisons between these two models did not indicate statistically significant differences in any of the three metrics. Random forest had a higher AUPRC and lower Brier score than gradient-boosted trees. Calibration-in-the-large was close to 0 for all three models, whereas the calibration slope was compatible with 1 for logistic regression and random forest but was below 1 for gradient-boosted trees. Knee swelling, limited range of motion, and injury mechanism were the most influential random-forest predictors. Conclusions: Repeated internal validation did not identify a uniformly superior model. Logistic regression most often ranked first for AUROC, whereas random forest most often ranked first for AUPRC and Brier score; however, paired comparisons did not identify significant differences between these two models. External validation is required before clinical use.

MedicinaVol. 62(10)
Daegu Catholic University (KR), Daegu Catholic University Medical Center (KR)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Total Knee Arthroplasty Outcomes
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.