SpecEdge: Adaptive Speculative Decoding and Quantization Dynamics for Edge Small Language Models
Auto-regressive sequence generation in modern transformer language models is severely memory-bandwidth bound, yielding low arithmetic intensity on resource-constrained edge devices. While speculative decoding mitigates this bottleneck by utilizing a compact draft model to propose candidates verified concurrently by a target model, its empirical behavior under aggressive low-bit post-training quantization remains largely unexplored. In this paper, we present SpecEdge, a systematic empirical investigation into the token acceptance dynamics, memory footprint, and wall-clock latency tradeoffs of speculative decoding across unquantized (FP16), 8-bit integer (INT8), and 4-bit NormalFloat (INT4-NF4) regimes on Small Language Models (SLMs, 135M–3B parameters). We demonstrate that low-bit quantization introduces significant probability distribution divergence, reducing static speculative acceptance rates by up to 21.8%. To resolve this divergence, we propose Entropy-Gated Adaptive Speculation (EG-Spec), a dynamic lookahead policy that monitors token-level Shannon entropy and prediction margins to terminate drafting prior to catastrophic verification rejections. On rigorous empirical benchmarks spanning mathematical reasoning (GSM8K), code generation (HumanEval), and abstractive summarization, SpecEdge achieves a 1.82×–2.38× wall-clock speedup over standard auto-regressive baselines and reduces peak VRAM by 58.2%, while mathematically maintaining exact target output distributions. Non-parametric hypothesis testing (paired Wilcoxon signed-rank test, p < 0.001) validates that EG-Spec consistently outperforms static lookahead horizons on quantized edge architectures.
Authors
- Benedict Baah
Institutions
- Jain University (IN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23186562
- Primary Topic
- Natural Language Processing Techniques
- Type
- preprint