An Efficient Transformer Detector for Traffic Scenes via Lightweight Backbone and Multi-Scale Attention

As autonomous driving systems shift from experimental testing environments to real-world open-road deployment, object detection in traffic scenes has emerged as a critical challenge in the perception pipeline. In complex traffic scenarios, significant variations in object scales (e.g., small vehicles and large infrastructure) and dynamic backgrounds often result in frequent missed detections and false positives under real-time constraints, which directly impairs driving safety. To tackle these issues, we propose Trans-DETR, an efficient end-to-end detector for traffic scenes based on a multi-scale attention mechanism and a lightweight backbone network. The conventional Convolutional Block Attention Module (CBAM) captures features at only a single scale, making it insufficient for handling multi-scale characteristics. To overcome this limitation, we introduce the Multi-Scale Attention Module (MSAM), which incorporates Multi-Scale Convolution Attention (MSCA) into the cross-scale fusion path of the model. This enables the effective extraction of multi-scale feature information and significantly improves detection accuracy. Furthermore, we propose the lightweight backbone network LiteSNet to reduce the parameter inflation and computational burden caused by traditional large-kernel convolutions. LiteSNet decouples large kernels into sequentially stacked horizontal and vertical separable convolutions, while employing dilated convolutions to expand the receptive field without information loss. This strategy significantly reduces both the parameter count and computational complexity while preserving the same effective receptive field as large kernels. Finally, we design a reparameterized detection head called RepHead, which deeply integrates cross-stage partial connections with reparameterization techniques. This allows for efficient fusion of low-level spatial details and high-level semantic features, thereby enhancing the model’s ability to detect small-scale traffic targets. Experimental results demonstrate that the proposed method achieves a mean average precision (mAP) of 89.8% on the KITTI dataset, outperforming RT-DETR-L by 2.0%. It also reduces the model size by 57% and computational complexity by 55.7%, with a single-frame inference time of only 4.6 ms. Compared with state-of-the-art approaches, the proposed algorithm achieves superior performance in terms of accuracy, efficiency, and model compactness, providing strong technical support for the advancement of autonomous driving systems.

Authors

Institutions

Publication Details

Journal
Sensors
Published
2026-09-24
DOI
https://doi.org/10.3390/s26196043
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

An Efficient Transformer Detector for Traffic Scenes via Lightweight Backbone and Multi-Scale Attention

Yanbo Hui, Xiaoxuan He, Xiaohui Zhang, Bo Chen et al.
Sensors
Advanced Neural Network Applications
article

An Efficient Transformer Detector for Traffic Scenes via Lightweight Backbone and Multi-Scale Attention

Yanbo Hui, Xiaoxuan He, Xiaohui Zhang, Bo Chen, Haiyang Ding, Yuan Zhang, Jinfeng Zheng, Zongpu Nan, Qiao Wang
article en

Abstract

As autonomous driving systems shift from experimental testing environments to real-world open-road deployment, object detection in traffic scenes has emerged as a critical challenge in the perception pipeline. In complex traffic scenarios, significant variations in object scales (e.g., small vehicles and large infrastructure) and dynamic backgrounds often result in frequent missed detections and false positives under real-time constraints, which directly impairs driving safety. To tackle these issues, we propose Trans-DETR, an efficient end-to-end detector for traffic scenes based on a multi-scale attention mechanism and a lightweight backbone network. The conventional Convolutional Block Attention Module (CBAM) captures features at only a single scale, making it insufficient for handling multi-scale characteristics. To overcome this limitation, we introduce the Multi-Scale Attention Module (MSAM), which incorporates Multi-Scale Convolution Attention (MSCA) into the cross-scale fusion path of the model. This enables the effective extraction of multi-scale feature information and significantly improves detection accuracy. Furthermore, we propose the lightweight backbone network LiteSNet to reduce the parameter inflation and computational burden caused by traditional large-kernel convolutions. LiteSNet decouples large kernels into sequentially stacked horizontal and vertical separable convolutions, while employing dilated convolutions to expand the receptive field without information loss. This strategy significantly reduces both the parameter count and computational complexity while preserving the same effective receptive field as large kernels. Finally, we design a reparameterized detection head called RepHead, which deeply integrates cross-stage partial connections with reparameterization techniques. This allows for efficient fusion of low-level spatial details and high-level semantic features, thereby enhancing the model’s ability to detect small-scale traffic targets. Experimental results demonstrate that the proposed method achieves a mean average precision (mAP) of 89.8% on the KITTI dataset, outperforming RT-DETR-L by 2.0%. It also reduces the model size by 57% and computational complexity by 55.7%, with a single-frame inference time of only 4.6 ms. Compared with state-of-the-art approaches, the proposed algorithm achieves superior performance in terms of accuracy, efficiency, and model compactness, providing strong technical support for the advancement of autonomous driving systems.

SensorsVol. 26(19)
Henan University of Technology (CN)
Industry, innovation and infrastructure
Openalex Percentile: Top 14%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.