Comparative Evaluation of Deep Learning and Hybrid CNN–Random Forest Models for Multi-Class Defect Identification in Railway Tracks
Railway infrastructure is fundamental to national logistics and public transportation systems. Defects in railway tracks, such as squats, shelling, spalling, flaking, burned rails, and joint issues, can significantly compromise operational safety. Traditional inspection methods, which rely heavily on manual labour, are often inefficient, error-prone, and infeasible for large-scale deployment. This study proposes a comparative experimental evaluation of deep learning models, and a hybrid CNN-feature/Random Forest model, for classifying surface-level defects in railway tracks from pre-cropped images. The dataset, comprising 6465 images (4525 training/1940 test) across six defect classes, was compiled from real-world conditions on the Indian railway network and used to train and evaluate five transfer-learning CNN backbones (MobileNetV2, DenseNet121, VGG16, ResNet50, EfficientNetB0) and a Random Forest classifier operating on CNN-extracted features. MobileNetV2 achieved the highest test-set accuracy (81.2%, F1 = 0.81, AUC-ROC = 0.96), with VGG16 and DenseNet121 close behind (80.7% and 80.0%); all three show a widening train/validation loss gap consistent with overfitting despite reasonable test-set generalisation. This aggregate accuracy conceals a safety-relevant weakness: MobileNetV2’s recall for Burned Rail, a safety-critical defect category, is only 45.5%, the lowest recall of any class for this model. ResNet50 and EfficientNetB0 underperformed substantially (54.2% and 42.0%); training logs and learning curves confirm this reflects a training/optimisation failure rather than a demonstrated architectural limitation; EfficientNetB0 in particular shows chance-level ROC-AUC (0.50) on every class, indicating its predictions carry no discriminative signal. A Random Forest classifier operating on VGG16 features achieved 72.2% accuracy on an independently partitioned test set, with similarly low recall for minority classes. The study discusses deployment trade-offs between accuracy and computational efficiency and their connection to existing non-destructive testing workflows, and identifies verification work that remains outstanding for the Random Forest hyperparameters and evaluation split.
Authors
- Ravikant Mordia (ORCID: https://orcid.org/0000-0003-3254-7724)
- Arvind Verma (ORCID: https://orcid.org/0000-0003-4177-4312)
Institutions
- Jodhpur National University (IN)
- Jai Narain Vyas University (IN)
Publication Details
- Journal
- NDT
- Published
- 2026-09-13
- DOI
- https://doi.org/10.3390/ndt4030027
- Primary Topic
- Railway Engineering and Dynamics
- Type
- article
- Field-Weighted Citation Impact
- 0.00