StrongerSORT: Improving DeepSORT for Stronger Human Tracking

Human tracking plays a crucial role in video surveillance systems. However, tracking humans in surveillance videos remains challenging because targets are often captured at long distances, occupy only a small number of pixels, and exhibit substantial scale variations. These challenges require not only accurate detection of small, low-texture targets but also fast and robust data association for multi-object tracking. We propose an enhanced human-tracking method that integrates improved object detection with complementary appearance and motion cues. Specifically, we improve YOLOv12 by incorporating large-kernel deformable attention, dynamic convolution, and phantom convolution. These components enhance the detector’s ability to perceive small-target features under complex backgrounds and occlusion while reducing its computational cost. The OSNet appearance embeddings incorporated into the EMA update framework construct a robust trajectory-level temporal appearance representation, which improves the discrimination between different individuals in the tracking stage. Extensive experiments demonstrate that the proposed method achieves more accurate identity association and more stable target trajectories than state-of-the-art tracking methods, including StrongSORT and ByteTrack.

Authors

Institutions

Publication Details

Journal
Sensors
Published
2026-09-17
DOI
https://doi.org/10.3390/s26185878
Primary Topic
Video Surveillance and Tracking Methods
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

StrongerSORT: Improving DeepSORT for Stronger Human Tracking

Jintao Sheng, Yali Zheng, Yinuo Wang, Jiayi Guan et al.
Sensors
Video Surveillance and Tracking Methods
article

StrongerSORT: Improving DeepSORT for Stronger Human Tracking

Jintao Sheng, Yali Zheng, Yinuo Wang, Jiayi Guan, Xinlu Zhong, Yunhua Tan, Da Lv
article en

Abstract

Human tracking plays a crucial role in video surveillance systems. However, tracking humans in surveillance videos remains challenging because targets are often captured at long distances, occupy only a small number of pixels, and exhibit substantial scale variations. These challenges require not only accurate detection of small, low-texture targets but also fast and robust data association for multi-object tracking. We propose an enhanced human-tracking method that integrates improved object detection with complementary appearance and motion cues. Specifically, we improve YOLOv12 by incorporating large-kernel deformable attention, dynamic convolution, and phantom convolution. These components enhance the detector’s ability to perceive small-target features under complex backgrounds and occlusion while reducing its computational cost. The OSNet appearance embeddings incorporated into the EMA update framework construct a robust trajectory-level temporal appearance representation, which improves the discrimination between different individuals in the tracking stage. Extensive experiments demonstrate that the proposed method achieves more accurate identity association and more stable target trajectories than state-of-the-art tracking methods, including StrongSORT and ByteTrack.

SensorsVol. 26(18)
University of Electronic Science and Technology of China (CN), Dongfang Electric Corporation (China) (CN), Institute for Advanced Study (DE)
National Natural Science Foundation of China
Reduced inequalities
Openalex Percentile: Top 14%
Video Surveillance and Tracking Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

StrongerSORT: Improving DeepSORT for Stronger Human Tracking — Jintao Sheng, Yali Zheng, et al. · Sensors (2026) | TGRS Research Map | TGRS