Dhanunjaya Varma Devalraju, Arshdeep Singh, Mark D. Plumbley · arXiv (Cornell University) 2026 · 2026
DOI: 10.5281/zenodo.23084982
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Natural language descriptions can provide rich semantic representations of audio-visual urban scenes, yet datasets that jointly describe both auditory and visual information remain limited. So, we introduce AVSD-Scenes, a paired audio-visual scene description dataset for urban environments. The dataset contains 12,291 audio-visual scene descriptions generated from the TAU Urban Audio-Visual Scenes dataset [1]. The audio-visual dataset examples [1] can be downloaded from https://zenodo.org/records/4477542 To construct the dataset, we first generate audio- and visual-based descriptions using Qwen2-Audio-7B and Qwen2.5-VL-7B, respectively. These modality specific descriptions are then combined using large language models, namely Qwen3-14B, Mistral-Small-3.2-24B-Instruct-2506, and Gemma-3-27B-it. Please visit our Github repository for codes, paper and more information. There are two files containing descriptions of the TAU Urban Audio-Visual Scenes dataset [1]. (1) master_captions xls : Descriptions when scene label is given as an instruction. (2) master_captions_without_scenelabel: Descriptions when scene label is not given as an instruction. The fields in the above files are: scene_label filename_audio audio_caption filename_video visual_caption multimodal_qwen3_caption multimodal_mistral_caption multimodal_gemma_caption where audio_caption: audio descriptions using Qwen2-Audio-7B, visual_caption: visual descriptions using Qwen2.5-VL-7B, multimodal_model_caption: audio-visual descriptions using models (Qwen3-14B, Mistral-Small-3.2-24B-Instruct-2506, and Gemma-3-27B-it.). [1] Wang, Shanshan, et al. "A curated dataset of urban scenes for audio-visual scene analysis." IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021. (Paper link)
No comments yet — start the discussion below.