Jialu Li, Jinchuan Tian, Shinji Watanabe · arXiv (Cornell University) 2026 · 2026
DOI: 10.48550/arxiv.2609.23114
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Recent advances in Speech Language Models (SpeechLMs), which integrate large language models with speech foundation models, have enabled unified sequence modeling of speech processing tasks. However, many SpeechLM-based approaches to speaker diarization (SD) are tightly coupled with automatic speech recognition (ASR) and evaluated using word-level metrics, making it difficult to assess SD performance independent of ASR accuracy. In this work, we investigate ESPnet-SpeechLM as a token-based backbone for generating SD hypotheses, formulating SD as autoregressive generation of structured tokens conditioned on acoustic input. We systematically compare two output representations: an event-based representation that explicitly models speaker turn onset and offset timestamps, and a frame-based representation that predicts frame-level speaker activity. To provide structured conversational cues, we further incorporate auxiliary tasks including speech activity detection, overlapped speech detection, and speaker turn counting within the output sequence. Across multiple meeting datasets, we find that event-based representations produce more stable and consistent SD outputs than frame-based representations. Our analysis shows that outputs generated by SpeechLMs encode useful temporal SD structure, but full-meeting SD remains limited by recording-level speaker tracking and overlap-related misses. Explicit speaker-linking post-processing substantially reduces speaker confusion, suggesting that robust SpeechLM-based SD requires persistent speaker tracking and overlap-aware generation.
No comments yet — start the discussion below.