Latest Research in Computer Vision and Pattern Recognition

107 research papers · 0.0 average citations · 2026 median publication year

Top Research Topics in Computer Vision and Pattern Recognition

Highest-Cited Papers

  1. Event-Grounded Football News Generation from Match Videos with Parameter-Efficient Large Language Models
  2. ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search
  3. SVMemAgent: A Streaming Video Memory Agent for Query-Agnostic Online Frame Selection
  4. Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
  5. AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction
  6. VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs
  7. CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video
  8. VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding
  9. Long-to-Short Video Evidence Reasoning for Grounded Question Answering
  10. MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding
  11. EventVL: Understand Event Streams via Multimodal Large Language Model
  12. Zero-shot video highlight detection based on text descriptions and synthetic images
  13. SGWIB:Sliced Gromov-Wasserstein Information Bottleneck for Video Highlight Detection
  14. Online Video Agent Harness for Long Video Understanding
  15. One Skill Does Not Fit All: Automatic Discovery and Taxonomy-Guided Routing of Frame-Selection Skills for Long-Video Question Answering
  16. ProactiveBench: Can Streaming Video Models Really Interact Like Humans?
  17. Routing Before Looking: Query-Adaptive Evidence Acquisition for Long-form Video Understanding
  18. StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs
  19. Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding
  20. Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
L3 Region - - 2026 Sep Q3

Computer Vision and Pattern Recognition

107 papers

Top Topics (10)

Computer Vision and Pattern Recognition85
Artificial Intelligence5
Intelligence, Security, War Strategy4
Multimodal Machine Learning Applications3
Information Retrieval2
Legal and Regulatory Analysis2
Video Analysis and Summarization1
Multimedia1
Machine Learning1
Public Relations and Crisis Communication1

Top Publications (20)

1.Event-Grounded Football News Generation from Match Videos with Parameter-Efficient Large Language Models2.ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search3.SVMemAgent: A Streaming Video Memory Agent for Query-Agnostic Online Frame Selection4.Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding5.AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction6.VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs7.CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video8.VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding9.Long-to-Short Video Evidence Reasoning for Grounded Question Answering10.MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding11.EventVL: Understand Event Streams via Multimodal Large Language Model12.Zero-shot video highlight detection based on text descriptions and synthetic images13.SGWIB:Sliced Gromov-Wasserstein Information Bottleneck for Video Highlight Detection14.Online Video Agent Harness for Long Video Understanding15.One Skill Does Not Fit All: Automatic Discovery and Taxonomy-Guided Routing of Frame-Selection Skills for Long-Video Question Answering16.ProactiveBench: Can Streaming Video Models Really Interact Like Humans?17.Routing Before Looking: Query-Adaptive Evidence Acquisition for Long-form Video Understanding18.StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs19.Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding20.Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.