Machine Learning-Predicted Length of Hospital Stay in Pediatric Appendicitis: An Exploratory Comparative Study of Clinical and Ultrasound-Based Models with Explainable Artificial Intelligence and Bayesian Analysis

Background: Accurate prediction of hospital length of stay (LOS) in pediatric appendicitis is essential for optimizing resource allocation, guiding discharge planning, and improving patient outcomes. This study aimed to explore the potential of machine learning (ML) models for predicting LOS using clinical and ultrasound features, compare the added value of ultrasound findings, and provide interpretable insights using explainable artificial intelligence (XAI) and Bayesian analysis. Methods: A retrospective cohort of 323 pediatric patients with uncomplicated appendicitis was analyzed from publicly available open-access data from the UCI Machine Learning Repository. Clinical-only and clinical-plus-ultrasound feature sets were used to train Random Forest, XGBoost, and Bayesian Additive Regression Trees (BART) models. Performance was evaluated using mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE), and P20 (percentage of predictions within 20% of actual values) with 95% confidence intervals (CIs) from bootstrap resampling. XAI analysis was employed for model interpretability, with exploratory, hypothesis-generating subgroup analyses and sensitivity analyses conducted across patient demographics and imputation strategies. Results: The Random Forest clinical model achieved the best performance (RMSE: 1.66 days, 95% CI: 1.20–2.15; MAE: 1.08 days, 95% CI: 0.80–1.42; P20: 58.7%, 95% CI: 46.0–69.8), significantly outperforming XGBoost (p = 0.03 for RMSE). BART demonstrated comparable accuracy (RMSE: 1.57 days, 95% CI: 1.21–2.12). SHAP analysis identified absence of peritonitis and appendix diameter as the most influential predictors. Subgroup analysis revealed a potentially better performance in female patients (RMSE: 1.21 days, P20: 75.8%) but poor accuracy for prolonged LOS ≥6 days (RMSE: 4.13 days, P20: 0%). Sensitivity analysis confirmed robustness of imputation (Spearman’s rho: 0.782–0.927). Conclusions: ML models, particularly Random Forest, predict LOS in pediatric appendicitis using clinical features. XAI provides clinically interpretable insights.

Authors

Institutions

Publication Details

Journal
Medical Sciences
Published
2026-09-09
DOI
https://doi.org/10.3390/medsci14050556
Primary Topic
Appendicitis Diagnosis and Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Machine Learning-Predicted Length of Hospital Stay in Pediatric Appendicitis: An Exploratory Comparative Study of Clinical and Ultrasound-Based Models with Explainable Artificial Intelligence and Bayesian Analysis

Kannan Sridharan, Gowri Sivaramakrishnan
Medical Sciences
Appendicitis Diagnosis and Management
article

Machine Learning-Predicted Length of Hospital Stay in Pediatric Appendicitis: An Exploratory Comparative Study of Clinical and Ultrasound-Based Models with Explainable Artificial Intelligence and Bayesian Analysis

Kannan Sridharan, Gowri Sivaramakrishnan
article en

Abstract

Background: Accurate prediction of hospital length of stay (LOS) in pediatric appendicitis is essential for optimizing resource allocation, guiding discharge planning, and improving patient outcomes. This study aimed to explore the potential of machine learning (ML) models for predicting LOS using clinical and ultrasound features, compare the added value of ultrasound findings, and provide interpretable insights using explainable artificial intelligence (XAI) and Bayesian analysis. Methods: A retrospective cohort of 323 pediatric patients with uncomplicated appendicitis was analyzed from publicly available open-access data from the UCI Machine Learning Repository. Clinical-only and clinical-plus-ultrasound feature sets were used to train Random Forest, XGBoost, and Bayesian Additive Regression Trees (BART) models. Performance was evaluated using mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE), and P20 (percentage of predictions within 20% of actual values) with 95% confidence intervals (CIs) from bootstrap resampling. XAI analysis was employed for model interpretability, with exploratory, hypothesis-generating subgroup analyses and sensitivity analyses conducted across patient demographics and imputation strategies. Results: The Random Forest clinical model achieved the best performance (RMSE: 1.66 days, 95% CI: 1.20–2.15; MAE: 1.08 days, 95% CI: 0.80–1.42; P20: 58.7%, 95% CI: 46.0–69.8), significantly outperforming XGBoost (p = 0.03 for RMSE). BART demonstrated comparable accuracy (RMSE: 1.57 days, 95% CI: 1.21–2.12). SHAP analysis identified absence of peritonitis and appendix diameter as the most influential predictors. Subgroup analysis revealed a potentially better performance in female patients (RMSE: 1.21 days, P20: 75.8%) but poor accuracy for prolonged LOS ≥6 days (RMSE: 4.13 days, P20: 0%). Sensitivity analysis confirmed robustness of imputation (Spearman’s rho: 0.782–0.927). Conclusions: ML models, particularly Random Forest, predict LOS in pediatric appendicitis using clinical features. XAI provides clinically interpretable insights.

Medical SciencesVol. 14(5)
Arabian Gulf University (BH)
Openalex Percentile: Top 7%
Appendicitis Diagnosis and Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.