An Efficient Transformer Detector for Traffic Scenes via Lightweight Backbone and Multi-Scale Attention
As autonomous driving systems shift from experimental testing environments to real-world open-road deployment, object detection in traffic scenes has emerged as a critical challenge in the perception pipeline. In complex traffic scenarios, significant variations in object scales (e.g., small vehicles and large infrastructure) and dynamic backgrounds often result in frequent missed detections and false positives under real-time constraints, which directly impairs driving safety. To tackle these issues, we propose Trans-DETR, an efficient end-to-end detector for traffic scenes based on a multi-scale attention mechanism and a lightweight backbone network. The conventional Convolutional Block Attention Module (CBAM) captures features at only a single scale, making it insufficient for handling multi-scale characteristics. To overcome this limitation, we introduce the Multi-Scale Attention Module (MSAM), which incorporates Multi-Scale Convolution Attention (MSCA) into the cross-scale fusion path of the model. This enables the effective extraction of multi-scale feature information and significantly improves detection accuracy. Furthermore, we propose the lightweight backbone network LiteSNet to reduce the parameter inflation and computational burden caused by traditional large-kernel convolutions. LiteSNet decouples large kernels into sequentially stacked horizontal and vertical separable convolutions, while employing dilated convolutions to expand the receptive field without information loss. This strategy significantly reduces both the parameter count and computational complexity while preserving the same effective receptive field as large kernels. Finally, we design a reparameterized detection head called RepHead, which deeply integrates cross-stage partial connections with reparameterization techniques. This allows for efficient fusion of low-level spatial details and high-level semantic features, thereby enhancing the model’s ability to detect small-scale traffic targets. Experimental results demonstrate that the proposed method achieves a mean average precision (mAP) of 89.8% on the KITTI dataset, outperforming RT-DETR-L by 2.0%. It also reduces the model size by 57% and computational complexity by 55.7%, with a single-frame inference time of only 4.6 ms. Compared with state-of-the-art approaches, the proposed algorithm achieves superior performance in terms of accuracy, efficiency, and model compactness, providing strong technical support for the advancement of autonomous driving systems.
Authors
- Yanbo Hui (ORCID: https://orcid.org/0009-0006-6743-5027)
- Xiaoxuan He (ORCID: https://orcid.org/0000-0002-2685-5010)
- Xiaohui Zhang (ORCID: https://orcid.org/0000-0001-9610-854X)
- Bo Chen (ORCID: https://orcid.org/0000-0002-9926-4377)
- Haiyang Ding (ORCID: https://orcid.org/0009-0003-6204-7525)
- Yuan Zhang
- Jinfeng Zheng (ORCID: https://orcid.org/0009-0005-3407-4176)
- Zongpu Nan
- Qiao Wang
Institutions
- Henan University of Technology (CN)
Publication Details
- Journal
- Sensors
- Published
- 2026-09-24
- DOI
- https://doi.org/10.3390/s26196043
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00