Streaming Video Moment Retrieval via Temporal Forecast Diffusion Model

With the advancement of online media, the demand for real-time interaction with streaming video has rapidly increased, highlighting the importance of Streaming Video Moment Retrieval (SVMR). This task aims to evaluate the relevance between streaming video and textual queries in real-time. Unlike offline videos, streaming videos must be processed on-the-fly, which presents substantial challenges due to the absence of future frames and the redundancy of historical frames. These factors lead to insufficient contextual information, which hinders reasoning about the current frame. To address these challenges, we propose DiffuStream, a novel approach from a generative perspective. Specifically, DiffuStream leverages a Temporal Forecast Diffusion Model to efficiently synthesize future semantics conditioned on observable contexts, thereby providing additional future information to assist in reasoning about the current frame. Complementarily, we introduce a Segment-based Dual Compressor to extract multi-granularity cues from noisy history while serving as a robust condition to assist future generation. By integrating the synthesized future with the refined history, DiffuStream constructs a comprehensive historical-present-future context, facilitating precise cross-modal reasoning. Extensive experiments on the ActivityNet Captions, TACoS, and MAD datasets validate the superior effectiveness of DiffuStream.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Multimedia Computing Communications and Applications
Published
2026-10-03
DOI
https://doi.org/10.1145/3849706
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Streaming Video Moment Retrieval via Temporal Forecast Diffusion Model

Cheng Dan Deng, Zhe Xu, Yanhua Yang, Kun Wei et al.
ACM Transactions on Multimedia Computing Communications and Applications
Multimodal Machine Learning Applications
article

Streaming Video Moment Retrieval via Temporal Forecast Diffusion Model

Cheng Dan Deng, Zhe Xu, Yanhua Yang, Kun Wei, Jiahua Li
article en

Abstract

With the advancement of online media, the demand for real-time interaction with streaming video has rapidly increased, highlighting the importance of Streaming Video Moment Retrieval (SVMR). This task aims to evaluate the relevance between streaming video and textual queries in real-time. Unlike offline videos, streaming videos must be processed on-the-fly, which presents substantial challenges due to the absence of future frames and the redundancy of historical frames. These factors lead to insufficient contextual information, which hinders reasoning about the current frame. To address these challenges, we propose DiffuStream, a novel approach from a generative perspective. Specifically, DiffuStream leverages a Temporal Forecast Diffusion Model to efficiently synthesize future semantics conditioned on observable contexts, thereby providing additional future information to assist in reasoning about the current frame. Complementarily, we introduce a Segment-based Dual Compressor to extract multi-granularity cues from noisy history while serving as a robust condition to assist future generation. By integrating the synthesized future with the refined history, DiffuStream constructs a comprehensive historical-present-future context, facilitating precise cross-modal reasoning. Extensive experiments on the ActivityNet Captions, TACoS, and MAD datasets validate the superior effectiveness of DiffuStream.

ACM Transactions on Multimedia Computing Communications and Applications
Xidian University (CN), Hong Kong University of Science and Technology (HK)
Openalex Percentile: Top 14%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Streaming Video Moment Retrieval via Temporal Forecast Diffusion Model — Cheng Dan Deng, Zhe Xu, et al. · ACM Transactions on Multimedia Computing Communications and Applications (2026) | TGRS Research Map | TGRS