Predicting rule-based dairy milk intake recommendations using machine learning: A robustness, calibration, and explainability study

Abstract Background Personalized dairy guidance is clinically complex because nutritional needs, intolerance, allergy, bone health, and cardiometabolic factors can point toward different decisions. Machine learning can help study how structured variables map to recommendations, but performance on synthetic labels must not be confused with clinical validity. Methods Five classifiers—Logistic Regression, Random Forest, LightGBM, K-Nearest Neighbors, and a soft-voting ensemble—were evaluated on the cited public synthetic dataset. An audit confirmed 10,000 records, 15 source predictors, one engineered age-group predictor, and a highly imbalanced three-class target: Maintain 7,981 (79.81%), Increase 2,002 (20.02%), and Reduce 17 (0.17%). The primary analysis used leakage-controlled stratified 10-fold cross-validation. Results Random Forest achieved the highest mean accuracy (0.9981 ± 0.0007) and macro ROC-AUC (1.0000 ± 0.0000), but its mean macro F1-score was 0.6661 ± 0.0005 and its geometric mean was 0 because it did not identify any of the 17 Reduce records in out-of-fold prediction. Logistic Regression achieved lower accuracy (0.9811 ± 0.0039) but the strongest balanced discrimination, with macro F1-score 0.8403 ± 0.0360, macro recall 0.9589 ± 0.0700, and geometric mean 0.9512 ± 0.0861. The overall paired difference in fold-level macro F1-score was significant (Friedman χ² = 26.88, p = 2.10 × 10 − 5 ), although most differences among LightGBM, Logistic Regression, and the voting ensemble were not significant after Holm correction. Conclusion The study provides a reproducible framework for interrogating machine learning behavior on rule-generated dairy recommendation labels. It does not establish a clinically validated recommendation system.

Authors

Institutions

Publication Details

Journal
Journal of Umm Al-Qura University for Medical Sciences
Published
2026-09-25
DOI
https://doi.org/10.1007/s44361-026-00058-w
Primary Topic
Nutritional Studies and Diet
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Predicting rule-based dairy milk intake recommendations using machine learning: A robustness, calibration, and explainability study

Hamza Shahbaz, Fazal Khaliq, Abdurrahman Sarnikli
Journal of Umm Al-Qura University for Medical Sciences
Nutritional Studies and Diet
article

Predicting rule-based dairy milk intake recommendations using machine learning: A robustness, calibration, and explainability study

Hamza Shahbaz, Fazal Khaliq, Abdurrahman Sarnikli
article en

Abstract

Abstract Background Personalized dairy guidance is clinically complex because nutritional needs, intolerance, allergy, bone health, and cardiometabolic factors can point toward different decisions. Machine learning can help study how structured variables map to recommendations, but performance on synthetic labels must not be confused with clinical validity. Methods Five classifiers—Logistic Regression, Random Forest, LightGBM, K-Nearest Neighbors, and a soft-voting ensemble—were evaluated on the cited public synthetic dataset. An audit confirmed 10,000 records, 15 source predictors, one engineered age-group predictor, and a highly imbalanced three-class target: Maintain 7,981 (79.81%), Increase 2,002 (20.02%), and Reduce 17 (0.17%). The primary analysis used leakage-controlled stratified 10-fold cross-validation. Results Random Forest achieved the highest mean accuracy (0.9981 ± 0.0007) and macro ROC-AUC (1.0000 ± 0.0000), but its mean macro F1-score was 0.6661 ± 0.0005 and its geometric mean was 0 because it did not identify any of the 17 Reduce records in out-of-fold prediction. Logistic Regression achieved lower accuracy (0.9811 ± 0.0039) but the strongest balanced discrimination, with macro F1-score 0.8403 ± 0.0360, macro recall 0.9589 ± 0.0700, and geometric mean 0.9512 ± 0.0861. The overall paired difference in fold-level macro F1-score was significant (Friedman χ² = 26.88, p = 2.10 × 10 − 5 ), although most differences among LightGBM, Logistic Regression, and the voting ensemble were not significant after Holm correction. Conclusion The study provides a reproducible framework for interrogating machine learning behavior on rule-generated dairy recommendation labels. It does not establish a clinically validated recommendation system.

Journal of Umm Al-Qura University for Medical SciencesVol. 12(2)
Aksaray University (TR), Government College University, Faisalabad (PK), The University of Agriculture, Peshawar (PK)
Peace, Justice and strong institutions
Openalex Percentile: Top 9%
Nutritional Studies and Diet
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.