Text Query-Guided Video Summarization with Self-Attentive Fully Convolutional Network

The importance of text queries in video summarization has grown significantly as they enable the generation of summaries tailored to individual user requirements. The summarization process in both conventional and domain-specific video summarization methods does not take into account the user’s preferences. Given the perspective nature of video summarization, it becomes crucial to incorporate user preferences while generating the video summaries. This research addresses this challenge by proposing a Self-Attentive Fully Convolutional Network (SAFCN), which takes an input text query from the user and generates a video summary pertinent to the query. The proposed network comprises a video representation unit utilizing a pretrained 3D Convolutional Neural Network (CNN), local self-attention, and query-relevant global self-attention units, which compute the importance scores of video shots based on the input textual query. The model incorporates a 2D positional encoding within the self-attention mechanisms and utilizes the Lion optimizer, along with summary length regularization. Extensive experimentation is conducted on the benchmark Query-Focused Video Summarization (QFVS) and QVHighlights dataset. The proposed method achieves a 5% improvement in F1-score for the QFVS dataset and demonstrates competitive performance on the QVHighlights dataset. Both qualitative and quantitative results illustrate that the proposed approach performs better than state-of-the-art techniques.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Asian and Low-Resource Language Information Processing
Published
2026-10-05
DOI
https://doi.org/10.1145/3856799
Primary Topic
Video Analysis and Summarization
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Text Query-Guided Video Summarization with Self-Attentive Fully Convolutional Network

Bhakti D. Kadam, Ashwini Mangesh Deshpande
ACM Transactions on Asian and Low-Resource Language Information Processing
Video Analysis and Summarization
article

Text Query-Guided Video Summarization with Self-Attentive Fully Convolutional Network

Bhakti D. Kadam, Ashwini Mangesh Deshpande
article en

Abstract

The importance of text queries in video summarization has grown significantly as they enable the generation of summaries tailored to individual user requirements. The summarization process in both conventional and domain-specific video summarization methods does not take into account the user’s preferences. Given the perspective nature of video summarization, it becomes crucial to incorporate user preferences while generating the video summaries. This research addresses this challenge by proposing a Self-Attentive Fully Convolutional Network (SAFCN), which takes an input text query from the user and generates a video summary pertinent to the query. The proposed network comprises a video representation unit utilizing a pretrained 3D Convolutional Neural Network (CNN), local self-attention, and query-relevant global self-attention units, which compute the importance scores of video shots based on the input textual query. The model incorporates a 2D positional encoding within the self-attention mechanisms and utilizes the Lion optimizer, along with summary length regularization. Extensive experimentation is conducted on the benchmark Query-Focused Video Summarization (QFVS) and QVHighlights dataset. The proposed method achieves a 5% improvement in F1-score for the QFVS dataset and demonstrates competitive performance on the QVHighlights dataset. Both qualitative and quantitative results illustrate that the proposed approach performs better than state-of-the-art techniques.

ACM Transactions on Asian and Low-Resource Language Information Processing
Defence Institute of Advanced Technology (IN)
Openalex Percentile: Top 14%
Video Analysis and Summarization
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Text Query-Guided Video Summarization with Self-Attentive Fully Convolutional Network — Bhakti D. Kadam, Ashwini Mangesh Deshpande · ACM Transactions on Asian and Low-Resource Language Information Processing (2026) | TGRS Research Map | TGRS