SISEVIR: From Manual Inspection to Automated Diagnosis of Vertical Traffic Signs Through YOLO Segmentation, EfficientNet, and Vision–Language Models for National Road Safety Management in Peru

Inventory and condition assessment of vertical traffic signs constitutes an essential activity for road safety management, as these signs serve as the primary mechanism for regulating vehicular flow and alerting drivers to prevailing road conditions. However, current inspection procedures in Peru rely on field crews that evaluate each sign manually, thereby constraining the frequency, objectivity, and scalability of the process. This paper presents SISEVIR (Sistema de Supervisión de Señales Verticales en Infraestructura Vial), a three-stage deep learning pipeline for the automated diagnosis of vertical traffic sign condition. The first stage employs YOLO26s-seg for instance segmentation of 31 sign classes, achieving a test mAP50 of 0.9305 (box) and 0.9224 (mask). The second stage classifies each detected sign into seven deterioration states using EfficientNet-B0, optimized through a five-experiment ablation study that identified progressive offline augmentation as the most effective strategy for handling a 147:1 class imbalance (macro F1 = 0.8252; pairwise McNemar’s tests with Holm–Bonferroni correction did not confirm significance at the family-wise α=0.05 level). The third stage integrates Qwen2-VL-2B-Instruct, a vision–language model, to generate natural-language descriptions of sign condition aligned with the MTC Manual of Traffic Control Devices for Streets and Highways. A structured evaluation by two independent raters on 35 descriptions yielded a correctness rate of 93.5% among valid responses (95% CI: 79.3–98.2%, Cohen’s κ=1.00). The system was trained and validated on a proprietary dataset of 5935 images and 6412 labeled crops collected along three routes in the Ayacucho Region (246.6 km total), with an inter-rater reliability of κ=0.802 (95% CI: 0.676–0.928). SISEVIR processes vehicular video at 30.7 FPS on an NVIDIA RTX 5080 GPU and assigns each sign a level within a four-tier condition scale (Optimal through Critical) linked to specific maintenance interventions, significantly reducing the time, cost, and personnel required compared with the manual inspection method established in the MSV-2016 Road Safety Manual.

Authors

Institutions

Publication Details

Journal
Future Transportation
Published
2026-08-28
DOI
https://doi.org/10.3390/futuretransp6050184
Primary Topic
Traffic and Road Safety
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

SISEVIR: From Manual Inspection to Automated Diagnosis of Vertical Traffic Signs Through YOLO Segmentation, EfficientNet, and Vision–Language Models for National Road Safety Management in Peru

Wilmer Moncada, Renato Soca Flores, Marco Castillo, Santiago Fernández et al.
Future Transportation
Traffic and Road Safety
article

SISEVIR: From Manual Inspection to Automated Diagnosis of Vertical Traffic Signs Through YOLO Segmentation, EfficientNet, and Vision–Language Models for National Road Safety Management in Peru

Wilmer Moncada, Renato Soca Flores, Marco Castillo, Santiago Fernández, Edwin Portal-Quicaña, Cristhian Aldana, Manuel Lagos, Yesenia Saavedra, Christian Lezama-Cuellar, Hemerson Lizarbe Alarcón, Rocky Giban Ayala Bizarro, Diego O. Tenorio-Huarancca, Victor Portal Quicaña, Kely Pilar Huaman de la Cruz
article en

Abstract

Inventory and condition assessment of vertical traffic signs constitutes an essential activity for road safety management, as these signs serve as the primary mechanism for regulating vehicular flow and alerting drivers to prevailing road conditions. However, current inspection procedures in Peru rely on field crews that evaluate each sign manually, thereby constraining the frequency, objectivity, and scalability of the process. This paper presents SISEVIR (Sistema de Supervisión de Señales Verticales en Infraestructura Vial), a three-stage deep learning pipeline for the automated diagnosis of vertical traffic sign condition. The first stage employs YOLO26s-seg for instance segmentation of 31 sign classes, achieving a test mAP50 of 0.9305 (box) and 0.9224 (mask). The second stage classifies each detected sign into seven deterioration states using EfficientNet-B0, optimized through a five-experiment ablation study that identified progressive offline augmentation as the most effective strategy for handling a 147:1 class imbalance (macro F1 = 0.8252; pairwise McNemar’s tests with Holm–Bonferroni correction did not confirm significance at the family-wise α=0.05 level). The third stage integrates Qwen2-VL-2B-Instruct, a vision–language model, to generate natural-language descriptions of sign condition aligned with the MTC Manual of Traffic Control Devices for Streets and Highways. A structured evaluation by two independent raters on 35 descriptions yielded a correctness rate of 93.5% among valid responses (95% CI: 79.3–98.2%, Cohen’s κ=1.00). The system was trained and validated on a proprietary dataset of 5935 images and 6412 labeled crops collected along three routes in the Ayacucho Region (246.6 km total), with an inter-rater reliability of κ=0.802 (95% CI: 0.676–0.928). SISEVIR processes vehicular video at 30.7 FPS on an NVIDIA RTX 5080 GPU and assigns each sign a level within a four-tier condition scale (Optimal through Critical) linked to specific maintenance interventions, significantly reducing the time, cost, and personnel required compared with the manual inspection method established in the MSV-2016 Road Safety Manual.

Future TransportationVol. 6(5)
Universidad de La Frontera (CL), San Cristóbal of Huamanga University (PE)
Sustainable cities and communities
Openalex Percentile: Top 11%
Traffic and Road Safety
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.