Latest Research in Multimodal Video Perception

118 research papers · 2026 median publication year

Top Research Topics in Multimodal Video Perception

Highest-Cited Papers

  1. Event-Grounded Football News Generation from Match Videos with Parameter-Efficient Large Language Models
  2. EuroMillions Breakthrough Mining — 4165 Sources, 37272 Leads Extracted — E8 Intelligence Research
  3. AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models
  4. EuroMillions Breakthrough Mining — 4119 Sources, 36796 Leads Extracted — E8 Intelligence Research
  5. ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search
  6. SVMemAgent: A Streaming Video Memory Agent for Query-Agnostic Online Frame Selection
  7. Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
  8. Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment
  9. AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction
  10. VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs
  11. CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video
  12. VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding
  13. Long-to-Short Video Evidence Reasoning for Grounded Question Answering
  14. MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding
  15. Hypergraph-Regularized Gramian Volumes for Multimodal Retrieval
  16. EventVL: Understand Event Streams via Multimodal Large Language Model
  17. Zero-shot video highlight detection based on text descriptions and synthetic images
  18. SGWIB:Sliced Gromov-Wasserstein Information Bottleneck for Video Highlight Detection
  19. Online Video Agent Harness for Long Video Understanding
  20. One Skill Does Not Fit All: Automatic Discovery and Taxonomy-Guided Routing of Frame-Selection Skills for Long-Video Question Answering

Sub-Regions

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
L2 Region - - 2026 Sep Q3

Multimodal Video Perception

118 papers

Top Topics (10)

Computer Vision and Pattern Recognition92
Intelligence, Security, War Strategy7
Artificial Intelligence5
Multimodal Machine Learning Applications4
Video Analysis and Summarization2
Information Retrieval2
Multimedia1
Machine Learning1
Public Relations and Crisis Communication1
Databases1

Top Publications (20)

1.Event-Grounded Football News Generation from Match Videos with Parameter-Efficient Large Language Models2.EuroMillions Breakthrough Mining — 4165 Sources, 37272 Leads Extracted — E8 Intelligence Research3.AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models4.EuroMillions Breakthrough Mining — 4119 Sources, 36796 Leads Extracted — E8 Intelligence Research5.ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search6.SVMemAgent: A Streaming Video Memory Agent for Query-Agnostic Online Frame Selection7.Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding8.Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment9.AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction10.VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs11.CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video12.VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding13.Long-to-Short Video Evidence Reasoning for Grounded Question Answering14.MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding15.Hypergraph-Regularized Gramian Volumes for Multimodal Retrieval16.EventVL: Understand Event Streams via Multimodal Large Language Model17.Zero-shot video highlight detection based on text descriptions and synthetic images18.SGWIB:Sliced Gromov-Wasserstein Information Bottleneck for Video Highlight Detection19.Online Video Agent Harness for Long Video Understanding20.One Skill Does Not Fit All: Automatic Discovery and Taxonomy-Guided Routing of Frame-Selection Skills for Long-Video Question Answering

Sub-Regions (3)

AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.