EAS-Cap: Event-Aware Semantic-Enhanced Video Captioning

Video captioning faces significant challenges in balancing global semantic consistency with fine-grained syntactic accuracy, particularly in complex multi-event scenarios. Existing global alignment methods often suffer from semantic dilution, while fine-grained approaches based on Part-of-Speech tagging frequently encounter compositional failures. To bridge these gaps, we propose EAS-Cap (Event-Aware Semantic-Enhanced Captioning), a novel framework that redefines the fundamental learning unit as Atomic Events—indivisible Subject-Verb-Object triplets extracted via syntactic analysis. EAS-Cap decomposes the captioning task into three synergistic objectives: atomic event extraction, cross-modal semantic distillation, and global-local consistency optimization. Specifically, our Event Module employs learnable queries to distill textual priors into visual encoders, enhancing sensitivity to specific actions, while a Visual-to-Event constraint ensures local features remain consistent with global context. Extensive experiments on the MSVD and MSR-VTT datasets demonstrate that EAS-Cap achieves excellent performance, notably leading in CIDEr scores. These results confirm our method's superior capability in preserving syntactic integrity and accurately describing complex temporal dynamics without relying on noisy global signals or explicit grammatical constraints.

Authors

Institutions

Publication Details

Journal
International Journal of Pattern Recognition and Artificial Intelligence
Published
2026-09-30
DOI
https://doi.org/10.1142/s0218001426590421
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

EAS-Cap: Event-Aware Semantic-Enhanced Video Captioning

Bin Fang, Langping Wang
International Journal of Pattern Recognition and Artificial Intelligence
Multimodal Machine Learning Applications
article

EAS-Cap: Event-Aware Semantic-Enhanced Video Captioning

Bin Fang, Langping Wang
article en

Abstract

Video captioning faces significant challenges in balancing global semantic consistency with fine-grained syntactic accuracy, particularly in complex multi-event scenarios. Existing global alignment methods often suffer from semantic dilution, while fine-grained approaches based on Part-of-Speech tagging frequently encounter compositional failures. To bridge these gaps, we propose EAS-Cap (Event-Aware Semantic-Enhanced Captioning), a novel framework that redefines the fundamental learning unit as Atomic Events—indivisible Subject-Verb-Object triplets extracted via syntactic analysis. EAS-Cap decomposes the captioning task into three synergistic objectives: atomic event extraction, cross-modal semantic distillation, and global-local consistency optimization. Specifically, our Event Module employs learnable queries to distill textual priors into visual encoders, enhancing sensitivity to specific actions, while a Visual-to-Event constraint ensures local features remain consistent with global context. Extensive experiments on the MSVD and MSR-VTT datasets demonstrate that EAS-Cap achieves excellent performance, notably leading in CIDEr scores. These results confirm our method's superior capability in preserving syntactic integrity and accurately describing complex temporal dynamics without relying on noisy global signals or explicit grammatical constraints.

International Journal of Pattern Recognition and Artificial Intelligence
Twitter (United States) (US)
Quality Education
Openalex Percentile: Top 14%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

EAS-Cap: Event-Aware Semantic-Enhanced Video Captioning — Bin Fang, Langping Wang · International Journal of Pattern Recognition and Artificial Intelligence (2026) | TGRS Research Map | TGRS