View while moving: a unified multi-modal paradigm for efficient recognition in long-untrimmed videos
Abstract Recent studies have shown that reducing redundancy and noise in various dimensions (e.g., temporal, spatial) is effective for efficient video recognition. However, most existing efficient methods for video recognition utilize a two-stage paradigm for adaptive selection, which is often described as “preview-then-recognition”. In this paper, as opposed to performing the “preview” and “recognition” separately, we explore a new unified paradigm, “View while Moving”, named ViMo, for efficient recognition of long-untrimmed videos. The two phases of sampling and recognition, from coarse-grained to fine-grained, are integrated into a unified process, requiring only one-time viewing of the raw frame during inference. The spatiotemporal modeling follows a hierarchical structure, beginning with the capture of unit-level features and subsequently reasoning about higher-level video semantics. What’s more, not content with focusing solely on attention allocation within an individual modality, we extend ViMo to multi-modal scenarios based on game theory, named ViMo+. The new multi-modal ViMo+ is capable of modeling redundancy and noise not only within unimodality but also across multi-modalities. Extensive experiments on both long-untrimmed and short-trimmed videos show that our ViMo and ViMo+ achieve state-of-the-art accuracy while significantly reducing the inference cost.
Authors
- Yimei Huang (ORCID: https://orcid.org/0000-0002-2011-4844)
- Lanshan Zhang (ORCID: https://orcid.org/0000-0002-0674-7864)
- Wendong Wang (ORCID: https://orcid.org/0000-0002-9041-1721)
- Wufan Wang
- Jingyi Wang
- Ye Tian
- Mengyu Yang
- Xirong Que
Publication Details
- Journal
- Tsinghua Science & Technology
- Published
- 2026-09-21
- DOI
- https://doi.org/10.26599/tst.2026.9010083
- Primary Topic
- Human Pose and Action Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00