Development of a machine learning-based predictive model for feeding intolerance risk in preterm infants

Feeding intolerance (FI) is a common complication in premature infants, often resulting in delayed achievement of full enteral feeding and prolonged hospitalization. This study aimed to construct an interpretable risk-prediction model for FI using machine-learning algorithms based on routine electronic medical-record data. We prospectively collected clinical data for 420 preterm infants admitted to a tertiary hospital between June 2023 and August 2024. Time zero was defined as birth, the intended prediction time point was admission to the Neonatal Intensive Care Unit, and the prediction horizon was the first week of life. To avoid data leakage, the cohort was split 70:30 (stratified by outcome) before any further processing; missForest-style imputation, LASSO feature selection (10-fold CV, λ_min rule), and hyper-parameter tuning were performed within the training set only (9 predictors retained). Nine machine-learning algorithms were compared: support vector machine (SVM), random forest (RF), artificial neural network (ANN), logistic regression (LR), extreme gradient boosting (XGBoost), elastic net (EN), light gradient boosting machine (LightGBM), k-nearest neighbor (KNN), and decision tree (DT). Model performance was assessed by area under the receiver-operating characteristic curve (AUC), accuracy, recall, precision, F1 score (all with 95% bootstrap CIs from 2000 resamples), calibration intercept and slope, and Brier score. SHapley Additive exPlanations (SHAP; KernelExplainer with k-means background of 100 training samples, nsamples = 2000) were used for model interpretation. The incidence of FI was 43.6% (183/420). After LASSO selection, 9 predictors remained. On the held-out test set (n = 126, 55 events), the KNN model achieved the most balanced performance: AUC = 0.862 (95% CI 0.795–0.922), F1 = 0.733 (95% CI 0.621–0.820), Brier = 0.146 and calibration slope = 1.062 (95% CI 0.711–1.643). Internal validation by 100 bootstrap optimism correction showed an optimism-corrected AUC of 0.865 for KNN, indicating only modest overfitting. SHAP analysis identified gastric-content withdrawal, caffeine therapy, tube feeding, PICC insertion, antibiotic use, defecation interval, gastric injury, birth weight, and sex as the most influential contributors to the model-predicted risk. A leakage-free KNN-based FI risk-prediction model constructed from routine clinical data demonstrated good discrimination and acceptable calibration in internal validation. Given the single-center design, the limited sample size, and the uncertain temporal precedence of some predictors, the model is exploratory and is not ready for clinical implementation; prospective studies with a predefined prediction time point, standardized predictor collection, and external validation are required.

Authors

Institutions

Publication Details

Journal
Medicine
Published
2026-10-09
DOI
https://doi.org/10.1097/md.0000000000051021
Primary Topic
Infant Nutrition and Health
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Development of a machine learning-based predictive model for feeding intolerance risk in preterm infants

Yu Hong, Tan Yufei, Zeng Xiqiu, Li Yin et al.
Medicine
Infant Nutrition and Health
article

Development of a machine learning-based predictive model for feeding intolerance risk in preterm infants

Yu Hong, Tan Yufei, Zeng Xiqiu, Li Yin, Long Zhuo, Shen Liangrong
article en

Abstract

Feeding intolerance (FI) is a common complication in premature infants, often resulting in delayed achievement of full enteral feeding and prolonged hospitalization. This study aimed to construct an interpretable risk-prediction model for FI using machine-learning algorithms based on routine electronic medical-record data. We prospectively collected clinical data for 420 preterm infants admitted to a tertiary hospital between June 2023 and August 2024. Time zero was defined as birth, the intended prediction time point was admission to the Neonatal Intensive Care Unit, and the prediction horizon was the first week of life. To avoid data leakage, the cohort was split 70:30 (stratified by outcome) before any further processing; missForest-style imputation, LASSO feature selection (10-fold CV, λ_min rule), and hyper-parameter tuning were performed within the training set only (9 predictors retained). Nine machine-learning algorithms were compared: support vector machine (SVM), random forest (RF), artificial neural network (ANN), logistic regression (LR), extreme gradient boosting (XGBoost), elastic net (EN), light gradient boosting machine (LightGBM), k-nearest neighbor (KNN), and decision tree (DT). Model performance was assessed by area under the receiver-operating characteristic curve (AUC), accuracy, recall, precision, F1 score (all with 95% bootstrap CIs from 2000 resamples), calibration intercept and slope, and Brier score. SHapley Additive exPlanations (SHAP; KernelExplainer with k-means background of 100 training samples, nsamples = 2000) were used for model interpretation. The incidence of FI was 43.6% (183/420). After LASSO selection, 9 predictors remained. On the held-out test set (n = 126, 55 events), the KNN model achieved the most balanced performance: AUC = 0.862 (95% CI 0.795–0.922), F1 = 0.733 (95% CI 0.621–0.820), Brier = 0.146 and calibration slope = 1.062 (95% CI 0.711–1.643). Internal validation by 100 bootstrap optimism correction showed an optimism-corrected AUC of 0.865 for KNN, indicating only modest overfitting. SHAP analysis identified gastric-content withdrawal, caffeine therapy, tube feeding, PICC insertion, antibiotic use, defecation interval, gastric injury, birth weight, and sex as the most influential contributors to the model-predicted risk. A leakage-free KNN-based FI risk-prediction model constructed from routine clinical data demonstrated good discrimination and acceptable calibration in internal validation. Given the single-center design, the limited sample size, and the uncertain temporal precedence of some predictors, the model is exploratory and is not ready for clinical implementation; prospective studies with a predefined prediction time point, standardized predictor collection, and external validation are required.

MedicineVol. 105(41)
First Affiliated Hospital of Xi'an Jiaotong University (CN), Shenzhen Children's Hospital (CN)
Openalex Percentile: Top 13%
Infant Nutrition and Health
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.