Comparative Evaluation of Deep Learning and Hybrid CNN–Random Forest Models for Multi-Class Defect Identification in Railway Tracks

Railway infrastructure is fundamental to national logistics and public transportation systems. Defects in railway tracks, such as squats, shelling, spalling, flaking, burned rails, and joint issues, can significantly compromise operational safety. Traditional inspection methods, which rely heavily on manual labour, are often inefficient, error-prone, and infeasible for large-scale deployment. This study proposes a comparative experimental evaluation of deep learning models, and a hybrid CNN-feature/Random Forest model, for classifying surface-level defects in railway tracks from pre-cropped images. The dataset, comprising 6465 images (4525 training/1940 test) across six defect classes, was compiled from real-world conditions on the Indian railway network and used to train and evaluate five transfer-learning CNN backbones (MobileNetV2, DenseNet121, VGG16, ResNet50, EfficientNetB0) and a Random Forest classifier operating on CNN-extracted features. MobileNetV2 achieved the highest test-set accuracy (81.2%, F1 = 0.81, AUC-ROC = 0.96), with VGG16 and DenseNet121 close behind (80.7% and 80.0%); all three show a widening train/validation loss gap consistent with overfitting despite reasonable test-set generalisation. This aggregate accuracy conceals a safety-relevant weakness: MobileNetV2’s recall for Burned Rail, a safety-critical defect category, is only 45.5%, the lowest recall of any class for this model. ResNet50 and EfficientNetB0 underperformed substantially (54.2% and 42.0%); training logs and learning curves confirm this reflects a training/optimisation failure rather than a demonstrated architectural limitation; EfficientNetB0 in particular shows chance-level ROC-AUC (0.50) on every class, indicating its predictions carry no discriminative signal. A Random Forest classifier operating on VGG16 features achieved 72.2% accuracy on an independently partitioned test set, with similarly low recall for minority classes. The study discusses deployment trade-offs between accuracy and computational efficiency and their connection to existing non-destructive testing workflows, and identifies verification work that remains outstanding for the Random Forest hyperparameters and evaluation split.

Authors

Institutions

Publication Details

Journal
NDT
Published
2026-09-13
DOI
https://doi.org/10.3390/ndt4030027
Primary Topic
Railway Engineering and Dynamics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Comparative Evaluation of Deep Learning and Hybrid CNN–Random Forest Models for Multi-Class Defect Identification in Railway Tracks

Ravikant Mordia, Arvind Verma
NDT
Railway Engineering and Dynamics
article

Comparative Evaluation of Deep Learning and Hybrid CNN–Random Forest Models for Multi-Class Defect Identification in Railway Tracks

Ravikant Mordia, Arvind Verma
article en

Abstract

Railway infrastructure is fundamental to national logistics and public transportation systems. Defects in railway tracks, such as squats, shelling, spalling, flaking, burned rails, and joint issues, can significantly compromise operational safety. Traditional inspection methods, which rely heavily on manual labour, are often inefficient, error-prone, and infeasible for large-scale deployment. This study proposes a comparative experimental evaluation of deep learning models, and a hybrid CNN-feature/Random Forest model, for classifying surface-level defects in railway tracks from pre-cropped images. The dataset, comprising 6465 images (4525 training/1940 test) across six defect classes, was compiled from real-world conditions on the Indian railway network and used to train and evaluate five transfer-learning CNN backbones (MobileNetV2, DenseNet121, VGG16, ResNet50, EfficientNetB0) and a Random Forest classifier operating on CNN-extracted features. MobileNetV2 achieved the highest test-set accuracy (81.2%, F1 = 0.81, AUC-ROC = 0.96), with VGG16 and DenseNet121 close behind (80.7% and 80.0%); all three show a widening train/validation loss gap consistent with overfitting despite reasonable test-set generalisation. This aggregate accuracy conceals a safety-relevant weakness: MobileNetV2’s recall for Burned Rail, a safety-critical defect category, is only 45.5%, the lowest recall of any class for this model. ResNet50 and EfficientNetB0 underperformed substantially (54.2% and 42.0%); training logs and learning curves confirm this reflects a training/optimisation failure rather than a demonstrated architectural limitation; EfficientNetB0 in particular shows chance-level ROC-AUC (0.50) on every class, indicating its predictions carry no discriminative signal. A Random Forest classifier operating on VGG16 features achieved 72.2% accuracy on an independently partitioned test set, with similarly low recall for minority classes. The study discusses deployment trade-offs between accuracy and computational efficiency and their connection to existing non-destructive testing workflows, and identifies verification work that remains outstanding for the Random Forest hyperparameters and evaluation split.

NDTVol. 4(3)
Jodhpur National University (IN), Jai Narain Vyas University (IN)
Openalex Percentile: Top 20%
Railway Engineering and Dynamics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.