Object-state-based frame sampling for efficient video object detection

Object detection in video streams suffers from severe temporal redundancy, which inflates training cost and can over-represent static backgrounds. This paper presents an object-state-based data-curation framework for offline training-data selection in video object detection. The sampler uses ground-truth or teacher-generated pseudo boxes to initialize active segments, retain object births, and select additional frames when per-object bounding-box states change sufficiently according to an Intersection-over-Union (IoU) trigger. The selection rule provides a lightweight geometric proxy for preserving object-state events that are relevant to localization supervision while suppressing near-duplicate frames. On VisDrone-VID, using an IoU threshold of τ = 0.2 achieves mAP@50-95 0.130 versus 0.122 for Full while using fewer frames; under a stricter τ = 0.0 matched-budget setting, the method remains close to Full and achieves the highest mean mAP@50-95 among the compressed methods, with modest margins over Uniform and Random sampling. KITTI and Kgalagadi diagnostics further characterize the method under object-dense and background-dominant regimes: under an approximately matched-step KITTI setting, the proposed sampler achieves 0.3673 ± 0.0051 mAP@50-95, comparable to Full 0.3657 ± 0.0021 and higher than Uniform, Random, and Motion_BBox; in the high-redundancy Kgalagadi stream, it reaches 0.3878 ± 0.0160 compared with 0.3623 ± 0.0279 for Uniform and 0.3491 ± 0.0186 for Random. Object-event and detector-signal diagnostics show that selected frames preserve object births and large box-change events and have higher average gradient norms and losses than discarded or unselected-control frames. The proprietary CCTV case study further shows that class-aligned cross-domain fusion with MS COCO and BDD100K can complement sampled CCTV frames by adding external appearance diversity across shared classes while retaining the CCTV domain as the anchor source of supervision. Overall, the framework preserves detection accuracy while reducing redundant video frames, and cross-domain fusion separately adds appearance-diversity.

Authors

Institutions

Publication Details

Journal
Heliyon
Published
2026-09-15
DOI
https://doi.org/10.1016/j.heliyon.2026.e45426
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Object-state-based frame sampling for efficient video object detection

Sangmin Lee, Hwayong Jeong
Heliyon
Advanced Neural Network Applications
article

Object-state-based frame sampling for efficient video object detection

Sangmin Lee, Hwayong Jeong
article en

Abstract

Object detection in video streams suffers from severe temporal redundancy, which inflates training cost and can over-represent static backgrounds. This paper presents an object-state-based data-curation framework for offline training-data selection in video object detection. The sampler uses ground-truth or teacher-generated pseudo boxes to initialize active segments, retain object births, and select additional frames when per-object bounding-box states change sufficiently according to an Intersection-over-Union (IoU) trigger. The selection rule provides a lightweight geometric proxy for preserving object-state events that are relevant to localization supervision while suppressing near-duplicate frames. On VisDrone-VID, using an IoU threshold of τ = 0.2 achieves mAP@50-95 0.130 versus 0.122 for Full while using fewer frames; under a stricter τ = 0.0 matched-budget setting, the method remains close to Full and achieves the highest mean mAP@50-95 among the compressed methods, with modest margins over Uniform and Random sampling. KITTI and Kgalagadi diagnostics further characterize the method under object-dense and background-dominant regimes: under an approximately matched-step KITTI setting, the proposed sampler achieves 0.3673 ± 0.0051 mAP@50-95, comparable to Full 0.3657 ± 0.0021 and higher than Uniform, Random, and Motion_BBox; in the high-redundancy Kgalagadi stream, it reaches 0.3878 ± 0.0160 compared with 0.3623 ± 0.0279 for Uniform and 0.3491 ± 0.0186 for Random. Object-event and detector-signal diagnostics show that selected frames preserve object births and large box-change events and have higher average gradient norms and losses than discarded or unselected-control frames. The proprietary CCTV case study further shows that class-aligned cross-domain fusion with MS COCO and BDD100K can complement sampled CCTV frames by adding external appearance diversity across shared classes while retaining the CCTV domain as the anchor source of supervision. Overall, the framework preserves detection accuracy while reducing redundant video frames, and cross-domain fusion separately adds appearance-diversity.

HeliyonVol. 12(15)
Kwangwoon University (KR)
Openalex Percentile: Top 13%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Object-state-based frame sampling for efficient video object detection — Sangmin Lee, Hwayong Jeong · Heliyon (2026) | TGRS Research Map | TGRS