Object-state-based frame sampling for efficient video object detection
Object detection in video streams suffers from severe temporal redundancy, which inflates training cost and can over-represent static backgrounds. This paper presents an object-state-based data-curation framework for offline training-data selection in video object detection. The sampler uses ground-truth or teacher-generated pseudo boxes to initialize active segments, retain object births, and select additional frames when per-object bounding-box states change sufficiently according to an Intersection-over-Union (IoU) trigger. The selection rule provides a lightweight geometric proxy for preserving object-state events that are relevant to localization supervision while suppressing near-duplicate frames. On VisDrone-VID, using an IoU threshold of τ = 0.2 achieves mAP@50-95 0.130 versus 0.122 for Full while using fewer frames; under a stricter τ = 0.0 matched-budget setting, the method remains close to Full and achieves the highest mean mAP@50-95 among the compressed methods, with modest margins over Uniform and Random sampling. KITTI and Kgalagadi diagnostics further characterize the method under object-dense and background-dominant regimes: under an approximately matched-step KITTI setting, the proposed sampler achieves 0.3673 ± 0.0051 mAP@50-95, comparable to Full 0.3657 ± 0.0021 and higher than Uniform, Random, and Motion_BBox; in the high-redundancy Kgalagadi stream, it reaches 0.3878 ± 0.0160 compared with 0.3623 ± 0.0279 for Uniform and 0.3491 ± 0.0186 for Random. Object-event and detector-signal diagnostics show that selected frames preserve object births and large box-change events and have higher average gradient norms and losses than discarded or unselected-control frames. The proprietary CCTV case study further shows that class-aligned cross-domain fusion with MS COCO and BDD100K can complement sampled CCTV frames by adding external appearance diversity across shared classes while retaining the CCTV domain as the anchor source of supervision. Overall, the framework preserves detection accuracy while reducing redundant video frames, and cross-domain fusion separately adds appearance-diversity.
Authors
- Sangmin Lee (ORCID: https://orcid.org/0000-0002-5215-2546)
- Hwayong Jeong
Institutions
- Kwangwoon University (KR)
Publication Details
- Journal
- Heliyon
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1016/j.heliyon.2026.e45426
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00