Ziyue Jiang, Zhou Li Zhao · INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION 2026 · 2026
DOI: 10.1145/3776574.3831131
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Effectively conveying emotional, tonal, and prosodic styles in complex scenarios is essential for text-to-speech (TTS) systems. Current solutions typically follow two paradigms: 1) zero-shot TTS systems deriving voice variability from reference speech prompts, or 2) instructed TTS systems relying on manual textual descriptions to control stylistic delivery—both presenting fundamental limitations in flexibility and automation. To address this, we introduce Thought-TTS, a novel TTS framework that utilizes chain-of-thought reasoning to automatically derive detailed style descriptions for each utterance. Specifically: 1) we propose a chain-of-thought style reasoning TTS model incorporating intermediate reasoning steps to generate style descriptions fitting the current scenario; 2) we adopt an architecture decoupling semantic and acoustic information to optimize the language model’s chain-of-thought reasoning capability; 3) we create CoTSpeech, a new benchmark dataset with 2,500 hours of speech and 1.1M labels for training and evaluating chain-of-thought enhanced TTS systems. Experimental results show that Thought-TTS successfully leverages chain-of-thought reasoning to comprehend context cues and produce accurate style descriptions through progressive thinking. The synthesized speech exhibits strong alignment with CoT-derived style attributes, enabling novel applications in controllable speech generation. The demo page is available on https://thought-tts-demo.github.io/thought-tts-demo/.
No comments yet — start the discussion below.