A Comparison of Coating-Defect Detection Models for Wind Turbine Towers
Wind energy is becoming increasingly important in meeting energy security, grid reliability, and the growing electricity demand. The rapid increase in the number and prevalence of wind turbines has triggered the need for fast, cost-effective, reliable, and easy-to-implement inspection systems. Automated detection of coating defects is important for the maintenance and corrosion prevention of the turbines. In this context, this study aimed to compare and evaluate the accuracy and efficiency of deep learning models for detection of coating defects from images. In this context, the study evaluates five object detection models, RT-DETR-L, YOLOv8-M, YOLOv9-C, YOLO11-L, and YOLO26-L, on an open coating-defect dataset, with inclusion, pinhole, and scratch classes. Each model was trained with 3 random seeds for images of 640 × 640 and 1088 × 1088-pixel resolutions. At 1088 × 1088, YOLO26-L achieved the highest mean mAP50–95 (0.1142), while RT-DETR-L achieved the highest recall (0.5591) and F1 score (0.2217) at the operating confidence threshold of 0.25; RT-DETR-L was also the slowest model (59.8 ms/image). YOLOv8-M achieved a mean mAP50–95 of 0.0976 together with the lowest latency (34.2 ms/image) and lowest true batch-size-1 peak GPU-memory allocation (278 MB). Reducing the resolution to 640 × 640 lowered mAP50–95 by 25.4% for YOLOv8-M and 33.6% for RT-DETR-L while reducing latency and memory use. The results revealed a pronounced accuracy-efficiency trade-off rather than universal superiority of one model for coating defect detection for wind turbine towers. The dataset was acquired during the wind-tower painting process under controlled industrial imaging conditions, so the results characterise in-process coating inspection rather than the inspection of weathered, in-service turbines. Because only three random seeds were used, paired comparisons between individual models were not statistically conclusive (8 of 100 pairwise tests reached p < 0.05) and the reported rankings are descriptive. Applying operating thresholds—determined on the validation set for each model—to the independent test set altered the architectural ranking. At 1088 × 1088 resolution, YOLO26-L surpassed RT-DETR-L to achieve the highest mean F1 score, whereas at 640 × 640 resolution, RT-DETR-L’s ranking dropped significantly. The fact that the selected thresholds range from 0.076 to 0.456 indicates that the differences observed at the common threshold of 0.25 are sensitive to the choice of threshold. While the mAP50–95 decreased by approximately half in the 640 × 640 evaluation—where data leakage was eliminated—the 1088 × 1088 results were considered optimistic, as they were obtained based on the previous segmentation.
Authors
- Si̇nan Melih Ni̇gdeli (ORCID: https://orcid.org/0000-0002-9517-7313)
- Yaren Aydın (ORCID: https://orcid.org/0000-0002-5134-9822)
- Ümit Işıkdağ (ORCID: https://orcid.org/0000-0002-2660-0106)
- Gebrai̇l Bekdaş (ORCID: https://orcid.org/0000-0002-7327-9810)
Institutions
- Mimar Sinan Güzel Sanatlar Üniversitesi (TR)
- Istanbul University-Cerrahpaşa (TR)
- Istanbul University (TR)
Publication Details
- Journal
- Coatings
- Published
- 2026-09-29
- DOI
- https://doi.org/10.3390/coatings16101159
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00