Benchmarking VLM Scene Understanding: How Well Do Vision-Language Models Interpret Video Narratives?
This study benchmarks three vision-language models Gemini 1.5 Pro, LLaMA 3.1 70B (via Groq), and LLaVA-NeXT-Video 7B , on structured video scene annotation across six dimensions: subject identification, action description, emotional tone, spatial relationships, scene transitions, and lighting/atmosphere. Results show proprietary models lead on temporal tasks, but all models struggle with emotional tone and spatial reasoning.
Authors
- Aurangzaib Shehzad Awan (ORCID: https://orcid.org/0009-0004-8285-7311)
Institutions
- National University of Computer and Emerging Sciences (PK)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22964601
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- preprint