Kaijie Xiao, Gao Y, Wei Dong · Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies 2026 · 2026
DOI: 10.1145/3831640
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Video analytics is ubiquitous in modern society, and the emergence of Multimodal Large Language Models (MLLMs) has made its application even more extensive. A common way to support MLLM-based video analytics is to continuously stream video to the cloud and then extracts visual features and samples frames for analysis. However, this cloud-centric architecture can introduce high transmission overhead. Simultaneously, the fixed-rate or uniform sampling approaches employed in these systems will miss critical visual evidence relevant to the query, resulting in reduced accuracy.; AB@In this paper, we propose UniOVA , an edge-cloud collaborative framework for universal on-demand video analytics. First, query-aware feature retrieval improves accuracy and reduces transmission overhead by retrieving and transmitting only relevant ViT features from the edge. Second, Interest-Aware ViT reduces edge overhead through hierarchical token merging which compresses irrelevant visual data based on user interests. Our evaluation shows that UniOVA achieves comparable or higher accuracy than LongVA while reducing transmission overhead by 37.1%-95.8%. Compared with ChatCam, UniOVA also improves accuracy by 63.2%-193.8%. In addition, on Jetson Xavier NX, UniOVA achieves 7.14-11.43 FPS for feature extraction, and its end-to-end query latency ranges from 3.67-8.18s under different bandwidth settings. A real-world user study further demonstrates its potential for efficient and precise video analytics tailored to individual user needs.
No comments yet — start the discussion below.