Yuxuan Jiang · · 2026
DOI: 10.21437/dc.2026-2
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Text-to-audio generation has advanced rapidly, producing remarkably diverse and high-fidelity audio from free-form natural language descriptions.However, these models still offer only coarse control and cannot customize the generated audio as users intend.We therefore aim to build a holistic modeling paradigm for controllable audio generation, extending the control capability across timing, long-form, acoustic, and speech content.We further identify four key research questions spanning data construction, model architecture, inference efficiency, and multi-task co-optimization.To this end, we develop four works, FreeAudio, ControlAudio, AnyAudio, and FreeSonic, that progressively tackle these questions and achieve strong experimental results.In the future, we will develop a more holistic modeling paradigm for accurate and versatile controllable audio generation that resolves these research questions.
No comments yet — start the discussion below.