UniOVA: Universal On-demand Video Analytics with Edge-Cloud Collaborative Multimodal LLM

Video analytics is ubiquitous in modern society, and the emergence of Multimodal Large Language Models (MLLMs) has made its application even more extensive. A common way to support MLLM-based video analytics is to continuously stream video to the cloud and then extracts visual features and samples frames for analysis. However, this cloud-centric architecture can introduce high transmission overhead. Simultaneously, the fixed-rate or uniform sampling approaches employed in these systems will miss critical visual evidence relevant to the query, resulting in reduced accuracy.; AB@In this paper, we propose UniOVA , an edge-cloud collaborative framework for universal on-demand video analytics. First, query-aware feature retrieval improves accuracy and reduces transmission overhead by retrieving and transmitting only relevant ViT features from the edge. Second, Interest-Aware ViT reduces edge overhead through hierarchical token merging which compresses irrelevant visual data based on user interests. Our evaluation shows that UniOVA achieves comparable or higher accuracy than LongVA while reducing transmission overhead by 37.1%-95.8%. Compared with ChatCam, UniOVA also improves accuracy by 63.2%-193.8%. In addition, on Jetson Xavier NX, UniOVA achieves 7.14-11.43 FPS for feature extraction, and its end-to-end query latency ranges from 3.67-8.18s under different bandwidth settings. A real-world user study further demonstrates its potential for efficient and precise video analytics tailored to individual user needs.

Authors

Publication Details

Journal
Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies
Published
2026-09-30
DOI
https://doi.org/10.1145/3831640
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

UniOVA: Universal On-demand Video Analytics with Edge-Cloud Collaborative Multimodal LLM

Wei Dong, Gao Y, Kaijie Xiao
Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies
Multimodal Machine Learning Applications
article

UniOVA: Universal On-demand Video Analytics with Edge-Cloud Collaborative Multimodal LLM

Wei Dong, Gao Y, Kaijie Xiao
article en

Abstract

Video analytics is ubiquitous in modern society, and the emergence of Multimodal Large Language Models (MLLMs) has made its application even more extensive. A common way to support MLLM-based video analytics is to continuously stream video to the cloud and then extracts visual features and samples frames for analysis. However, this cloud-centric architecture can introduce high transmission overhead. Simultaneously, the fixed-rate or uniform sampling approaches employed in these systems will miss critical visual evidence relevant to the query, resulting in reduced accuracy.; AB@In this paper, we propose UniOVA , an edge-cloud collaborative framework for universal on-demand video analytics. First, query-aware feature retrieval improves accuracy and reduces transmission overhead by retrieving and transmitting only relevant ViT features from the edge. Second, Interest-Aware ViT reduces edge overhead through hierarchical token merging which compresses irrelevant visual data based on user interests. Our evaluation shows that UniOVA achieves comparable or higher accuracy than LongVA while reducing transmission overhead by 37.1%-95.8%. Compared with ChatCam, UniOVA also improves accuracy by 63.2%-193.8%. In addition, on Jetson Xavier NX, UniOVA achieves 7.14-11.43 FPS for feature extraction, and its end-to-end query latency ranges from 3.67-8.18s under different bandwidth settings. A real-world user study further demonstrates its potential for efficient and precise video analytics tailored to individual user needs.

Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous TechnologiesVol. 10(3)
Openalex Percentile: Top 14%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

UniOVA: Universal On-demand Video Analytics with Edge-Cloud Collaborative Multimodal LLM — Wei Dong, Gao Y, et al. · Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies (2026) | TGRS Research Map | TGRS