Aurangzaib Shehzad Awan · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22964601
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This study benchmarks three vision-language models Gemini 1.5 Pro, LLaMA 3.1 70B (via Groq), and LLaVA-NeXT-Video 7B , on structured video scene annotation across six dimensions: subject identification, action description, emotional tone, spatial relationships, scene transitions, and lighting/atmosphere. Results show proprietary models lead on temporal tasks, but all models struggle with emotional tone and spatial reasoning.
No comments yet — start the discussion below.