Latest Research in Computer Vision and Pattern Recognition
107 research papers · 0.0 average citations · 2026 median publication year
Top Research Topics in Computer Vision and Pattern Recognition
- Computer Vision and Pattern Recognition — 85 papers
- Artificial Intelligence — 5 papers
- Intelligence, Security, War Strategy — 4 papers
- Multimodal Machine Learning Applications — 3 papers
- Information Retrieval — 2 papers
- Legal and Regulatory Analysis — 2 papers
- Video Analysis and Summarization — 1 papers
- Multimedia — 1 papers
- Machine Learning — 1 papers
- Public Relations and Crisis Communication — 1 papers
Highest-Cited Papers
- Event-Grounded Football News Generation from Match Videos with Parameter-Efficient Large Language Models
- ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search
- SVMemAgent: A Streaming Video Memory Agent for Query-Agnostic Online Frame Selection
- Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
- AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction
- VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs
- CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video
- VideoScout: Learning Agentic Active Exploration with Adaptive Reasoning Pacing for Long Video Understanding
- Long-to-Short Video Evidence Reasoning for Grounded Question Answering
- MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding
- EventVL: Understand Event Streams via Multimodal Large Language Model
- Zero-shot video highlight detection based on text descriptions and synthetic images
- SGWIB:Sliced Gromov-Wasserstein Information Bottleneck for Video Highlight Detection
- Online Video Agent Harness for Long Video Understanding
- One Skill Does Not Fit All: Automatic Discovery and Taxonomy-Guided Routing of Frame-Selection Skills for Long-Video Question Answering
- ProactiveBench: Can Streaming Video Models Really Interact Like Humans?
- Routing Before Looking: Query-Adaptive Evidence Acquisition for Long-form Video Understanding
- StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs
- Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding
- Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding