StrongerSORT: Improving DeepSORT for Stronger Human Tracking
Human tracking plays a crucial role in video surveillance systems. However, tracking humans in surveillance videos remains challenging because targets are often captured at long distances, occupy only a small number of pixels, and exhibit substantial scale variations. These challenges require not only accurate detection of small, low-texture targets but also fast and robust data association for multi-object tracking. We propose an enhanced human-tracking method that integrates improved object detection with complementary appearance and motion cues. Specifically, we improve YOLOv12 by incorporating large-kernel deformable attention, dynamic convolution, and phantom convolution. These components enhance the detector’s ability to perceive small-target features under complex backgrounds and occlusion while reducing its computational cost. The OSNet appearance embeddings incorporated into the EMA update framework construct a robust trajectory-level temporal appearance representation, which improves the discrimination between different individuals in the tracking stage. Extensive experiments demonstrate that the proposed method achieves more accurate identity association and more stable target trajectories than state-of-the-art tracking methods, including StrongSORT and ByteTrack.
Authors
- Jintao Sheng (ORCID: https://orcid.org/0000-0002-8626-5741)
- Yali Zheng (ORCID: https://orcid.org/0000-0002-2906-7984)
- Yinuo Wang
- Jiayi Guan
- Xinlu Zhong
- Yunhua Tan
- Da Lv
Institutions
- University of Electronic Science and Technology of China (CN)
- Dongfang Electric Corporation (China) (CN)
- Institute for Advanced Study (DE)
Publication Details
- Journal
- Sensors
- Published
- 2026-09-17
- DOI
- https://doi.org/10.3390/s26185878
- Primary Topic
- Video Surveillance and Tracking Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Natural Science Foundation of China