Monte Carlo cross-validation driven benchmarking of CNN and transformer-based segmentation models for airfield sealed and unsealed cracks analysis

Recent applications of deep learning (DL) for automated crack detection and length quantification are important for pavement maintenance. However, most studies rely on a single training–validation split, limiting assessment of model robustness and performance variability. This study addresses this gap using Monte Carlo cross-validation (MC-CV) for segmentation and crack length quantification. The need for MC-CV-based robustness analysis was first demonstrated on the DeepCrack dataset, where U-Net showed ±1.06% variation in F1 score across fifteen replicates compared with previously reported single-split results. Five DL models, YOLOv8m, YOLOv9c, YOLOv11m, SegFormer-B1, and U-Net, were then evaluated on an in-house airfield dataset containing predominantly sealed cracks from Custer Airport (TTF) and unsealed cracks from Osage Municipal Airport (D02). An 85/15 training–validation split provided the best-balanced performance. Optimal confidence thresholds were 0.15 for YOLOv8m and 0.20 for YOLOv9c and YOLOv11m. YOLOv8m achieved mean intersection of union (IOU) and Dice scores of 59.5% and 73.5%, with crack-length errors of 15.5% and 8.8% for TTF and D02, respectively. Statistical reliability analysis showed that 98 MC-CV replicates achieved a ±0.5% margin of error at 95% confidence, with an average crack-length error of 11.6%. Overall, the findings highlight the importance of MC-CV for consistent evaluation of DL models for crack detection and length quantification, supporting its broader adoption in pavement inspection research.

Authors

Institutions

Publication Details

Journal
International Journal of Pavement Engineering
Published
2026-09-15
DOI
https://doi.org/10.1080/10298436.2026.2732149
Primary Topic
Infrastructure Maintenance and Monitoring
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Monte Carlo cross-validation driven benchmarking of CNN and transformer-based segmentation models for airfield sealed and unsealed cracks analysis

Rajrup Mitra, Daniel Offenbacker, Md. Abdullah All Sourav, In Ho Cho et al.
International Journal of Pavement Engineering
Infrastructure Maintenance and Monitoring
article

Monte Carlo cross-validation driven benchmarking of CNN and transformer-based segmentation models for airfield sealed and unsealed cracks analysis

Rajrup Mitra, Daniel Offenbacker, Md. Abdullah All Sourav, In Ho Cho, Wensheng Zhang, Berk Gülmezoğlu, Kyle Potvin, Yunjeong Mo, Beena Ajmera, Hali̇l Ceylan, Ayhan Öner Yücel, Sunghwan Kim
article en

Abstract

Recent applications of deep learning (DL) for automated crack detection and length quantification are important for pavement maintenance. However, most studies rely on a single training–validation split, limiting assessment of model robustness and performance variability. This study addresses this gap using Monte Carlo cross-validation (MC-CV) for segmentation and crack length quantification. The need for MC-CV-based robustness analysis was first demonstrated on the DeepCrack dataset, where U-Net showed ±1.06% variation in F1 score across fifteen replicates compared with previously reported single-split results. Five DL models, YOLOv8m, YOLOv9c, YOLOv11m, SegFormer-B1, and U-Net, were then evaluated on an in-house airfield dataset containing predominantly sealed cracks from Custer Airport (TTF) and unsealed cracks from Osage Municipal Airport (D02). An 85/15 training–validation split provided the best-balanced performance. Optimal confidence thresholds were 0.15 for YOLOv8m and 0.20 for YOLOv9c and YOLOv11m. YOLOv8m achieved mean intersection of union (IOU) and Dice scores of 59.5% and 73.5%, with crack-length errors of 15.5% and 8.8% for TTF and D02, respectively. Statistical reliability analysis showed that 98 MC-CV replicates achieved a ±0.5% margin of error at 95% confidence, with an average crack-length error of 11.6%. Overall, the findings highlight the importance of MC-CV for consistent evaluation of DL models for crack detection and length quantification, supporting its broader adoption in pavement inspection research.

International Journal of Pavement EngineeringVol. 27(1)
Federal Aviation Administration (US), Iowa State University (US), Applied Technologies (United States) (US), Tennessee Technological University (US), Adnan Menderes University (TR)
Federal Aviation Administration
Openalex Percentile: Top 18%
Infrastructure Maintenance and Monitoring
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.