Optimized hybrid machine learning for interpretable prediction of crack-healing percentage in self-healing concrete

Abstract Self-healing concrete can potentially repair cracks on its own; however, predicting its healing performance is challenging because of the complex interactions among mixture characteristics, crack properties, healing-agent content, and healing time. This research explores the use of machine learning methods to predict the crack-healing percentage of self-healing concrete and compares four modeling approaches: Linear Regression (LR), Decision Tree optimized with Grey Wolf Optimization (DT + GWO), Artificial Neural Network optimized with Grey Wolf Optimization (ANN + GWO), and Least Squares Boosting optimized with Grey Wolf Optimization (LSBoost + GWO). A dataset of 1271 measurements was gathered from six independent studies, using initial crack width (ICW), mixture proportions, healing-agent content, and healing time as inputs, and crack-healing percentage as the output. Model performance was assessed using R, MSE, RMSE, MAE, MAD, and the a20 index, and 10-fold cross-validation was used to evaluate the variability and robustness of the predictive performance. LSBoost + GWO showed the strongest overall predictive performance among the evaluated models, with R values of 0.810, 0.704, and 0.666 for the training, testing, and validation sets, respectively. In comparison, the testing R values for DT + GWO, ANN + GWO, and LR were 0.678, 0.683, and 0.572, respectively. The 10-fold cross-validation of LSBoost + GWO produced an average R of 0.641, indicating variability across the folds. To further examine the observed differences between the two optimized models, a paired Wilcoxon signed-rank test was conducted using the same cross-validation folds. The results revealed no statistically significant differences between DT + GWO and LSBoost + GWO for the main error- and correlation-based metrics. Although LSBoost + GWO achieved a higher a20 value than DT + GWO, this difference was not statistically significant after multiple-comparison correction (adjusted p = 0.334). SHAP analysis identified Healing Time, B/g, and ICW as the most influential input variables, while FA/C showed a moderate contribution, and CA/C and W/C exhibited relatively smaller global effects. Additionally, two-dimensional partial dependence analysis was used to explore interactions among selected input variables. Overall, the findings suggest that machine learning, especially LSBoost + GWO, can serve as a data-driven exploratory tool for estimating crack-healing percentage and investigating the relative importance and interactions of factors associated with the healing behavior of self-healing concrete.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-10-06
DOI
https://doi.org/10.1038/s41598-026-74710-1
Primary Topic
Microbial Applications in Construction Materials
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Optimized hybrid machine learning for interpretable prediction of crack-healing percentage in self-healing concrete

Mahdi Nematzadeh, Mohammad Bahram
Scientific Reports
Microbial Applications in Construction Materials
article

Optimized hybrid machine learning for interpretable prediction of crack-healing percentage in self-healing concrete

Mahdi Nematzadeh, Mohammad Bahram
article en

Abstract

Abstract Self-healing concrete can potentially repair cracks on its own; however, predicting its healing performance is challenging because of the complex interactions among mixture characteristics, crack properties, healing-agent content, and healing time. This research explores the use of machine learning methods to predict the crack-healing percentage of self-healing concrete and compares four modeling approaches: Linear Regression (LR), Decision Tree optimized with Grey Wolf Optimization (DT + GWO), Artificial Neural Network optimized with Grey Wolf Optimization (ANN + GWO), and Least Squares Boosting optimized with Grey Wolf Optimization (LSBoost + GWO). A dataset of 1271 measurements was gathered from six independent studies, using initial crack width (ICW), mixture proportions, healing-agent content, and healing time as inputs, and crack-healing percentage as the output. Model performance was assessed using R, MSE, RMSE, MAE, MAD, and the a20 index, and 10-fold cross-validation was used to evaluate the variability and robustness of the predictive performance. LSBoost + GWO showed the strongest overall predictive performance among the evaluated models, with R values of 0.810, 0.704, and 0.666 for the training, testing, and validation sets, respectively. In comparison, the testing R values for DT + GWO, ANN + GWO, and LR were 0.678, 0.683, and 0.572, respectively. The 10-fold cross-validation of LSBoost + GWO produced an average R of 0.641, indicating variability across the folds. To further examine the observed differences between the two optimized models, a paired Wilcoxon signed-rank test was conducted using the same cross-validation folds. The results revealed no statistically significant differences between DT + GWO and LSBoost + GWO for the main error- and correlation-based metrics. Although LSBoost + GWO achieved a higher a20 value than DT + GWO, this difference was not statistically significant after multiple-comparison correction (adjusted p = 0.334). SHAP analysis identified Healing Time, B/g, and ICW as the most influential input variables, while FA/C showed a moderate contribution, and CA/C and W/C exhibited relatively smaller global effects. Additionally, two-dimensional partial dependence analysis was used to explore interactions among selected input variables. Overall, the findings suggest that machine learning, especially LSBoost + GWO, can serve as a data-driven exploratory tool for estimating crack-healing percentage and investigating the relative importance and interactions of factors associated with the healing behavior of self-healing concrete.

Scientific Reports
Openalex Percentile: Top 19%
Microbial Applications in Construction Materials
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.