Automatic Road-Crack Detection with Self-Supervised YOLOv7 in Drone Imagery
Supervised object detection models rely on large training datasets to achieve good performance. However, data labeling is often an expensive and time-consuming task. To mitigate this challenge, self-supervised methods have been utilized to learn representations from unlabeled data. Accordingly, this study uses self-supervised learning to improve the YOLOv7 model. YOLOv7 was chosen because of its reliable design and lightweight E-ELAN backbone network, which helps it learn faster without influencing the gradient path. As such, YOLOv7 undergoes two separate training stages: pre-training and fine-tuning. During pre-training, new synthetic images are automatically generated by fusing different kinds of foreground road-damage objects with various background images. In addition, a contrastive loss function was incorporated into YOLOv7 to make class-specific instances cluster together. In the second stage, the pre-trained YOLOv7 model is fine-tuned using real training images. To validate the effectiveness of the proposed approach, it has been evaluated on drone images for road-damage detection from the UAPD dataset with six classes: transverse, longitudinal, oblique, pothole, alligator, and crack repair. The numerical findings showed that self-supervised learning improved YOLOv7’s performance by over eight percentage points (mAP 81.7% vs. 73.5%). The proposed approach also outperformed other detectors found in the literature, including Faster R-CNN (48.8% mAP), self-supervised DETR (20.10% mAP), YOLOv3 (68.75% mAP), and YOLOv4 (56.6% mAP).
Authors
- Hussein Samma (ORCID: https://orcid.org/0000-0002-3562-2788)
Institutions
- King Fahd University of Petroleum and Minerals (SA)
Publication Details
- Journal
- Automation
- Published
- 2026-09-28
- DOI
- https://doi.org/10.3390/automation7050155
- Primary Topic
- Infrastructure Maintenance and Monitoring
- Type
- article
- Field-Weighted Citation Impact
- 0.00