EAS-Cap: Event-Aware Semantic-Enhanced Video Captioning
Video captioning faces significant challenges in balancing global semantic consistency with fine-grained syntactic accuracy, particularly in complex multi-event scenarios. Existing global alignment methods often suffer from semantic dilution, while fine-grained approaches based on Part-of-Speech tagging frequently encounter compositional failures. To bridge these gaps, we propose EAS-Cap (Event-Aware Semantic-Enhanced Captioning), a novel framework that redefines the fundamental learning unit as Atomic Events—indivisible Subject-Verb-Object triplets extracted via syntactic analysis. EAS-Cap decomposes the captioning task into three synergistic objectives: atomic event extraction, cross-modal semantic distillation, and global-local consistency optimization. Specifically, our Event Module employs learnable queries to distill textual priors into visual encoders, enhancing sensitivity to specific actions, while a Visual-to-Event constraint ensures local features remain consistent with global context. Extensive experiments on the MSVD and MSR-VTT datasets demonstrate that EAS-Cap achieves excellent performance, notably leading in CIDEr scores. These results confirm our method's superior capability in preserving syntactic integrity and accurately describing complex temporal dynamics without relying on noisy global signals or explicit grammatical constraints.
Authors
- Bin Fang (ORCID: https://orcid.org/0000-0003-1955-6626)
- Langping Wang
Institutions
- Twitter (United States) (US)
Publication Details
- Journal
- International Journal of Pattern Recognition and Artificial Intelligence
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1142/s0218001426590421
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00