View while moving: a unified multi-modal paradigm for efficient recognition in long-untrimmed videos

Abstract Recent studies have shown that reducing redundancy and noise in various dimensions (e.g., temporal, spatial) is effective for efficient video recognition. However, most existing efficient methods for video recognition utilize a two-stage paradigm for adaptive selection, which is often described as “preview-then-recognition”. In this paper, as opposed to performing the “preview” and “recognition” separately, we explore a new unified paradigm, “View while Moving”, named ViMo, for efficient recognition of long-untrimmed videos. The two phases of sampling and recognition, from coarse-grained to fine-grained, are integrated into a unified process, requiring only one-time viewing of the raw frame during inference. The spatiotemporal modeling follows a hierarchical structure, beginning with the capture of unit-level features and subsequently reasoning about higher-level video semantics. What’s more, not content with focusing solely on attention allocation within an individual modality, we extend ViMo to multi-modal scenarios based on game theory, named ViMo+. The new multi-modal ViMo+ is capable of modeling redundancy and noise not only within unimodality but also across multi-modalities. Extensive experiments on both long-untrimmed and short-trimmed videos show that our ViMo and ViMo+ achieve state-of-the-art accuracy while significantly reducing the inference cost.

Authors

Publication Details

Journal
Tsinghua Science & Technology
Published
2026-09-21
DOI
https://doi.org/10.26599/tst.2026.9010083
Primary Topic
Human Pose and Action Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

View while moving: a unified multi-modal paradigm for efficient recognition in long-untrimmed videos

Yimei Huang, Lanshan Zhang, Wendong Wang, Wufan Wang et al.
Tsinghua Science & Technology
Human Pose and Action Recognition
article

View while moving: a unified multi-modal paradigm for efficient recognition in long-untrimmed videos

Yimei Huang, Lanshan Zhang, Wendong Wang, Wufan Wang, Jingyi Wang, Ye Tian, Mengyu Yang, Xirong Que
article en

Abstract

Abstract Recent studies have shown that reducing redundancy and noise in various dimensions (e.g., temporal, spatial) is effective for efficient video recognition. However, most existing efficient methods for video recognition utilize a two-stage paradigm for adaptive selection, which is often described as “preview-then-recognition”. In this paper, as opposed to performing the “preview” and “recognition” separately, we explore a new unified paradigm, “View while Moving”, named ViMo, for efficient recognition of long-untrimmed videos. The two phases of sampling and recognition, from coarse-grained to fine-grained, are integrated into a unified process, requiring only one-time viewing of the raw frame during inference. The spatiotemporal modeling follows a hierarchical structure, beginning with the capture of unit-level features and subsequently reasoning about higher-level video semantics. What’s more, not content with focusing solely on attention allocation within an individual modality, we extend ViMo to multi-modal scenarios based on game theory, named ViMo+. The new multi-modal ViMo+ is capable of modeling redundancy and noise not only within unimodality but also across multi-modalities. Extensive experiments on both long-untrimmed and short-trimmed videos show that our ViMo and ViMo+ achieve state-of-the-art accuracy while significantly reducing the inference cost.

Tsinghua Science & Technology
Openalex Percentile: Top 13%
Human Pose and Action Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

View while moving: a unified multi-modal paradigm for efficient recognition in long-untrimmed videos — Yimei Huang, Lanshan Zhang, et al. · Tsinghua Science & Technology (2026) | TGRS Research Map | TGRS