Towards Optimal Speculative Decoding: A Comprehensive Survey on Optimizations and Future Directions
Speculative decoding (SD) has emerged as a compelling approach to breaking the sequential bottleneck of autoregressive large language model (LLM) inference. By coupling lightweight drafting with parallel verification, SD reduces expensive target-model decoding while preserving generation fidelity. Yet its realized speedup often falls short of its theoretical potential, as inefficiencies persist throughout the speculative decoding pipeline. This survey presents a unified analysis of the factors that ultimately bound the acceleration of SD. We examine how draft effectiveness, verification efficiency, and system execution jointly determine end-to-end performance, and critically synthesize the trade-offs underlying existing optimization directions. This perspective clarifies where speculative gains are created, where they are lost, and why isolated improvements do not necessarily translate into practical speedup. We further identify the unresolved challenges that continue to limit SD and outline promising directions toward more efficient and scalable LLM inference.
Authors
- Xiaohong Qian (ORCID: https://orcid.org/0000-0002-9012-4351)
- Wenjian Xu (ORCID: https://orcid.org/0000-0001-8852-2601)
- Jian Wan (ORCID: https://orcid.org/0000-0001-9882-3029)
- Lei Zhang (ORCID: https://orcid.org/0000-0002-6400-1884)
- Yuyang Ji (ORCID: https://orcid.org/0009-0005-8483-9659)
Institutions
- Zhejiang University of Science and Technology (CN)
- Institute of Computing Technology (CN)
- Zhejiang Lab (CN)
- Zhejiang University of Water Resource and Electric Power (CN)
Publication Details
- Journal
- ACM Computing Surveys
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1145/3846171
- Primary Topic
- Generative Adversarial Networks and Image Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00