Mixture replication in benchmark datasets inflates reported machine learning performance and degrades prediction interval coverage for ultra-high-performance concrete

Abstract Machine learning models for ultra-high-performance concrete (UHPC) commonly report coefficients of determination above 0.94 on held-out records, and these values now shape expectations for data-driven mix design. We show that a substantial part of this reported predictive performance depends on how the models are evaluated. Benchmark UHPC datasets contain replicate records, meaning the same mixture tested at several curing ages or specimen sizes. Random train-test splitting distributes these replicates across both sides of the split, so models are scored partly on mixtures encountered during training. Using an open dataset of 1,193 UHPC records and an independent benchmark of 1,030 conventional concrete records, we compared random cross-validation with mixture-grouped cross-validation for eight regressors and four target properties. Under mixture-grouped evaluation the coefficient of determination for compressive strength fell from 0.947 to 0.893 and for flexural strength from 0.897 to 0.779, while root-mean-square error rose by 42 to 47% relative to the value implied by random splitting. A controlled experiment in which sample size was held constant and only the number of records per mixture varied reproduced the effect on both datasets, with the gap growing from approximately zero at one record per mixture to 0.13 at four. Conformal prediction intervals calibrated on randomly selected records covered new mixtures 82 to 84% of the time against a 90% nominal level, whereas calibration on held-out mixtures restored approximately nominal coverage with wider intervals. We provide an evaluation and reporting protocol for concrete machine learning.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-24
DOI
https://doi.org/10.1038/s41598-026-70561-y
Primary Topic
Innovative concrete reinforcement materials
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Mixture replication in benchmark datasets inflates reported machine learning performance and degrades prediction interval coverage for ultra-high-performance concrete

SM Abdul Khader, Huda Aldosari, Murtaza M. Junaid Farooque, Quadri Noorulhasan Naveed et al.
Scientific Reports
Innovative concrete reinforcement materials
article

Mixture replication in benchmark datasets inflates reported machine learning performance and degrades prediction interval coverage for ultra-high-performance concrete

SM Abdul Khader, Huda Aldosari, Murtaza M. Junaid Farooque, Quadri Noorulhasan Naveed, Naseema Shaik, Shazia Jaffari, Irfan Anjum Badruddin, Altamashuddin Khan Nadeemallah
article en

Abstract

Abstract Machine learning models for ultra-high-performance concrete (UHPC) commonly report coefficients of determination above 0.94 on held-out records, and these values now shape expectations for data-driven mix design. We show that a substantial part of this reported predictive performance depends on how the models are evaluated. Benchmark UHPC datasets contain replicate records, meaning the same mixture tested at several curing ages or specimen sizes. Random train-test splitting distributes these replicates across both sides of the split, so models are scored partly on mixtures encountered during training. Using an open dataset of 1,193 UHPC records and an independent benchmark of 1,030 conventional concrete records, we compared random cross-validation with mixture-grouped cross-validation for eight regressors and four target properties. Under mixture-grouped evaluation the coefficient of determination for compressive strength fell from 0.947 to 0.893 and for flexural strength from 0.897 to 0.779, while root-mean-square error rose by 42 to 47% relative to the value implied by random splitting. A controlled experiment in which sample size was held constant and only the number of records per mixture varied reproduced the effect on both datasets, with the gap growing from approximately zero at one record per mixture to 0.13 at four. Conformal prediction intervals calibrated on randomly selected records covered new mixtures 82 to 84% of the time against a 90% nominal level, whereas calibration on held-out mixtures restored approximately nominal coverage with wider intervals. We provide an evaluation and reporting protocol for concrete machine learning.

Scientific Reports
National Institute of Technology Karnataka (IN), Prince Sattam Bin Abdulaziz University (SA), Manipal Academy of Higher Education (IN), Nilai University (MY), International Islamic University Malaysia (MY), Dhofar University (OM), King Khalid University (SA)
Openalex Percentile: Top 17%
Innovative concrete reinforcement materials
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.