An overview of discrete diffusion models in natural language processing
Abstract Discrete Diffusion Models (DDMs) have recently emerged as a promising paradigm for text generation, providing an alternative to the dominant autoregressive (AR) approach. Unlike AR models that rely on sequential decoding and suffer from exposure bias, DDMs operate in discrete token space through iterative denoising, enabling parallel generation and bidirectional context modeling. This design substantially reduces inference latency and facilitates controllable, structure-aware synthesis. Recent studies demonstrate that large-scale DDMs achieve performance comparable to, and in some cases surpassing, similarly sized AR models, underscoring their potential as a new foundation for natural language processing. In this survey, we review the theoretical underpinnings of DDMs, trace their evolution from early prototypes to billion-parameter architectures, and examine key strategies for training and inference. We further highlight representative applications ranging from text infilling and editing to multimodal generation, and conclude with open challenges and future directions that may shape the role of DDMs in next-generation language technologies.
Authors
- Zhongjiang He (ORCID: https://orcid.org/0009-0000-1835-9271)
- Kaidong Yu
- Yixuan Li
- Yongxiang Li
- Shuangyong Song
Institutions
- China Telecom (China) (CN)
- China Telecom (CN)
- Xi'an Jiaotong University (CN)
Publication Details
- Journal
- Vicinagearth.
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1007/s44336-026-00040-5
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00