Streaming Video Moment Retrieval via Temporal Forecast Diffusion Model
With the advancement of online media, the demand for real-time interaction with streaming video has rapidly increased, highlighting the importance of Streaming Video Moment Retrieval (SVMR). This task aims to evaluate the relevance between streaming video and textual queries in real-time. Unlike offline videos, streaming videos must be processed on-the-fly, which presents substantial challenges due to the absence of future frames and the redundancy of historical frames. These factors lead to insufficient contextual information, which hinders reasoning about the current frame. To address these challenges, we propose DiffuStream, a novel approach from a generative perspective. Specifically, DiffuStream leverages a Temporal Forecast Diffusion Model to efficiently synthesize future semantics conditioned on observable contexts, thereby providing additional future information to assist in reasoning about the current frame. Complementarily, we introduce a Segment-based Dual Compressor to extract multi-granularity cues from noisy history while serving as a robust condition to assist future generation. By integrating the synthesized future with the refined history, DiffuStream constructs a comprehensive historical-present-future context, facilitating precise cross-modal reasoning. Extensive experiments on the ActivityNet Captions, TACoS, and MAD datasets validate the superior effectiveness of DiffuStream.
Authors
- Cheng Dan Deng (ORCID: https://orcid.org/0000-0003-2620-3247)
- Zhe Xu (ORCID: https://orcid.org/0000-0001-6898-3443)
- Yanhua Yang (ORCID: https://orcid.org/0000-0002-7916-3683)
- Kun Wei (ORCID: https://orcid.org/0000-0001-7228-5322)
- Jiahua Li (ORCID: https://orcid.org/0009-0006-4144-8815)
Institutions
- Xidian University (CN)
- Hong Kong University of Science and Technology (HK)
Publication Details
- Journal
- ACM Transactions on Multimedia Computing Communications and Applications
- Published
- 2026-10-03
- DOI
- https://doi.org/10.1145/3849706
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00