Yixuan Li, Yixuan Li, Kaidong Yu, Shuangyong Song, Yongxiang Li, Yongxiang Li, Zhongjiang He · Vicinagearth. 2026 · 2026
DOI: 10.1007/s44336-026-00040-5
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Discrete Diffusion Models (DDMs) have recently emerged as a promising paradigm for text generation, providing an alternative to the dominant autoregressive (AR) approach. Unlike AR models that rely on sequential decoding and suffer from exposure bias, DDMs operate in discrete token space through iterative denoising, enabling parallel generation and bidirectional context modeling. This design substantially reduces inference latency and facilitates controllable, structure-aware synthesis. Recent studies demonstrate that large-scale DDMs achieve performance comparable to, and in some cases surpassing, similarly sized AR models, underscoring their potential as a new foundation for natural language processing. In this survey, we review the theoretical underpinnings of DDMs, trace their evolution from early prototypes to billion-parameter architectures, and examine key strategies for training and inference. We further highlight representative applications ranging from text infilling and editing to multimodal generation, and conclude with open challenges and future directions that may shape the role of DDMs in next-generation language technologies.
No comments yet — start the discussion below.