Latest Research in Multimodal Video Perception
118 research papers · 2026 median publication year
Top Research Topics in Multimodal Video Perception
- Computer Vision and Pattern Recognition — 92 papers
- Intelligence, Security, War Strategy — 7 papers
- Artificial Intelligence — 5 papers
- Multimodal Machine Learning Applications — 4 papers
- Video Analysis and Summarization — 2 papers
- Information Retrieval — 2 papers
- Multimedia — 1 papers
- Machine Learning — 1 papers
- Public Relations and Crisis Communication — 1 papers
- Databases — 1 papers
Highest-Cited Papers
- Event-Grounded Football News Generation from Match Videos with Parameter-Efficient Large Language Models
- EuroMillions Breakthrough Mining — 4165 Sources, 37272 Leads Extracted — E8 Intelligence Research
- AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models
- EuroMillions Breakthrough Mining — 4119 Sources, 36796 Leads Extracted — E8 Intelligence Research
- ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search
- SVMemAgent: A Streaming Video Memory Agent for Query-Agnostic Online Frame Selection
- Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
- Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment
- AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction
- VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs
- CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video
- VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding
- Long-to-Short Video Evidence Reasoning for Grounded Question Answering
- MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding
- Hypergraph-Regularized Gramian Volumes for Multimodal Retrieval
- EventVL: Understand Event Streams via Multimodal Large Language Model
- Zero-shot video highlight detection based on text descriptions and synthetic images
- SGWIB:Sliced Gromov-Wasserstein Information Bottleneck for Video Highlight Detection
- Online Video Agent Harness for Long Video Understanding
- One Skill Does Not Fit All: Automatic Discovery and Taxonomy-Guided Routing of Frame-Selection Skills for Long-Video Question Answering
Sub-Regions
- Computer Vision and Pattern Recognition — 107 papers
- Intelligence, Security, War Strategy — 16 papers
- Multimodal Machine Learning Applications — 8 papers