Towards Optimal Speculative Decoding: A Comprehensive Survey on Optimizations and Future Directions

Speculative decoding (SD) has emerged as a compelling approach to breaking the sequential bottleneck of autoregressive large language model (LLM) inference. By coupling lightweight drafting with parallel verification, SD reduces expensive target-model decoding while preserving generation fidelity. Yet its realized speedup often falls short of its theoretical potential, as inefficiencies persist throughout the speculative decoding pipeline. This survey presents a unified analysis of the factors that ultimately bound the acceleration of SD. We examine how draft effectiveness, verification efficiency, and system execution jointly determine end-to-end performance, and critically synthesize the trade-offs underlying existing optimization directions. This perspective clarifies where speculative gains are created, where they are lost, and why isolated improvements do not necessarily translate into practical speedup. We further identify the unresolved challenges that continue to limit SD and outline promising directions toward more efficient and scalable LLM inference.

Authors

Institutions

Publication Details

Journal
ACM Computing Surveys
Published
2026-09-16
DOI
https://doi.org/10.1145/3846171
Primary Topic
Generative Adversarial Networks and Image Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Towards Optimal Speculative Decoding: A Comprehensive Survey on Optimizations and Future Directions

Xiaohong Qian, Wenjian Xu, Jian Wan, Lei Zhang et al.
ACM Computing Surveys
Generative Adversarial Networks and Image Synthesis
article

Towards Optimal Speculative Decoding: A Comprehensive Survey on Optimizations and Future Directions

Xiaohong Qian, Wenjian Xu, Jian Wan, Lei Zhang, Yuyang Ji
article en

Abstract

Speculative decoding (SD) has emerged as a compelling approach to breaking the sequential bottleneck of autoregressive large language model (LLM) inference. By coupling lightweight drafting with parallel verification, SD reduces expensive target-model decoding while preserving generation fidelity. Yet its realized speedup often falls short of its theoretical potential, as inefficiencies persist throughout the speculative decoding pipeline. This survey presents a unified analysis of the factors that ultimately bound the acceleration of SD. We examine how draft effectiveness, verification efficiency, and system execution jointly determine end-to-end performance, and critically synthesize the trade-offs underlying existing optimization directions. This perspective clarifies where speculative gains are created, where they are lost, and why isolated improvements do not necessarily translate into practical speedup. We further identify the unresolved challenges that continue to limit SD and outline promising directions toward more efficient and scalable LLM inference.

ACM Computing Surveys
Zhejiang University of Science and Technology (CN), Institute of Computing Technology (CN), Zhejiang Lab (CN), Zhejiang University of Water Resource and Electric Power (CN)
Openalex Percentile: Top 13%
Generative Adversarial Networks and Image Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Towards Optimal Speculative Decoding: A Comprehensive Survey on Optimizations and Future Directions — Xiaohong Qian, Wenjian Xu, et al. · ACM Computing Surveys (2026) | TGRS Research Map | TGRS