An Enhanced RT-DETR with Selective Boundary Aggregation and Tri-Focal CIoU Loss for Real-Time Object Detection
Real-Time Detection Transformer (RT-DETR) is the first detection transformer-based model to achieve real-time performance. However, the concatenation operation treats all scale features equally without adaptive weighting, leading to insufficient multi-scale feature fusion. In addition, Generalized Intersection over Union (GIoU) does not explicitly consider the center distance and aspect ratio between the predicted bounding boxes and the ground-truth boxes, leading to slow convergence. To address the above problems, this paper proposes Tri-Focal DEtection TRansformer (TriDETR) by replacing the GIoU loss function with the Tri-Focal Complete Intersection over Union (CIoU) loss to optimize overlap and assign weights, allowing the model to focus on difficult patterns. Additionally, TriDETR uses Selective Boundary Aggregation (SBA) to better fuse multi-scale features. Experiments were conducted on PASCAL VOC 2007, Foggy Cityscapes, and ACDC datasets to demonstrate the effectiveness of the proposed method. Specifically, in terms of mAP 50:95 , TriDETR outperforms RT-DETR with the same backbone by 0.80-0.98 on PASCAL VOC, 0.20-1.23 on Foggy Cityscapes, and 0.30-0.95 on ACDC (snow, rain, and nighttime). Additionally, TriDETR outperforms the best-performing YOLO26 variant while maintaining approximately the same frames per second.
Authors
- Phạm Ngọc Hùng (ORCID: https://orcid.org/0000-0002-5584-5823)
- Duc-Anh Nguyen (ORCID: https://orcid.org/0000-0002-6337-6254)
- Dinh-Nhat Loi
- Cong-Huu Hoang
- Duc-Hung Nguyen
Publication Details
- Journal
- International Journal of Pattern Recognition and Artificial Intelligence
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1142/s0218001426550189
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00