Hui Wang, Yunfei Chu, Xize Cheng, Yu Xi, Meng Gao, Qi Chen, Yifan Yang, Wenxiang Guo, Yong Qin, Jin Xu · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22857938
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Audio-language models increasingly process long recordings and multiple audio inputs, requiring fine-grained acoustic understanding, cross-audio reasoning, and decisions under goals and constraints. Existing benchmarks cover these requirements unevenly, while answer accuracy can obscure failures in evidence localization and decision reliability. We introduce LAMAR-Bench, a diagnostic benchmark with 1,597 questions across three capability families: fine-grained understanding, cross-audio relation reasoning, and decisions under goals and constraints. A shared-fact pipeline derives tasks with traceable supervision from podcast recordings. Across the dataset, input counts range from 1 to 19 audio clips per question, with total input duration reaching 59.9 minutes. The evaluation combines task-level metrics, textual reference baselines, and paired diagnostics of answer format, presentation order, and requested goals. Experiments with 14 audio-language and omni-modal systems reveal persistent limitations: the highest paralinguistic-analysis and acoustic-localization accuracies are only 51.09% and 54.38%, respectively, while audio difference analysis remains challenging. A matched answer-format diagnostic reveals substantial gaps between answer recognition and direct evidence localization. Among four hosted systems evaluated on paired goals, the best both-goal accuracy is 63.64%, indicating that correct selections do not consistently extend across different requirements for the same candidates. These findings highlight the need for more precise evidence grounding and reliable goal-conditioned audio reasoning.
No comments yet — start the discussion below.