SpecEdge: Adaptive Speculative Decoding and Quantization Dynamics for Edge Small Language Models

Auto-regressive sequence generation in modern transformer language models is severely memory-bandwidth bound, yielding low arithmetic intensity on resource-constrained edge devices. While speculative decoding mitigates this bottleneck by utilizing a compact draft model to propose candidates verified concurrently by a target model, its empirical behavior under aggressive low-bit post-training quantization remains largely unexplored. In this paper, we present SpecEdge, a systematic empirical investigation into the token acceptance dynamics, memory footprint, and wall-clock latency tradeoffs of speculative decoding across unquantized (FP16), 8-bit integer (INT8), and 4-bit NormalFloat (INT4-NF4) regimes on Small Language Models (SLMs, 135M–3B parameters). We demonstrate that low-bit quantization introduces significant probability distribution divergence, reducing static speculative acceptance rates by up to 21.8%. To resolve this divergence, we propose Entropy-Gated Adaptive Speculation (EG-Spec), a dynamic lookahead policy that monitors token-level Shannon entropy and prediction margins to terminate drafting prior to catastrophic verification rejections. On rigorous empirical benchmarks spanning mathematical reasoning (GSM8K), code generation (HumanEval), and abstractive summarization, SpecEdge achieves a 1.82×–2.38× wall-clock speedup over standard auto-regressive baselines and reduces peak VRAM by 58.2%, while mathematically maintaining exact target output distributions. Non-parametric hypothesis testing (paired Wilcoxon signed-rank test, p < 0.001) validates that EG-Spec consistently outperforms static lookahead horizons on quantized edge architectures.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-06
DOI
https://doi.org/10.5281/zenodo.23186562
Primary Topic
Natural Language Processing Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

SpecEdge: Adaptive Speculative Decoding and Quantization Dynamics for Edge Small Language Models

Benedict Baah
Zenodo (CERN European Organization for Nuclear Research)
Natural Language Processing Techniques
preprint

SpecEdge: Adaptive Speculative Decoding and Quantization Dynamics for Edge Small Language Models

Benedict Baah
preprint en

Abstract

Auto-regressive sequence generation in modern transformer language models is severely memory-bandwidth bound, yielding low arithmetic intensity on resource-constrained edge devices. While speculative decoding mitigates this bottleneck by utilizing a compact draft model to propose candidates verified concurrently by a target model, its empirical behavior under aggressive low-bit post-training quantization remains largely unexplored. In this paper, we present SpecEdge, a systematic empirical investigation into the token acceptance dynamics, memory footprint, and wall-clock latency tradeoffs of speculative decoding across unquantized (FP16), 8-bit integer (INT8), and 4-bit NormalFloat (INT4-NF4) regimes on Small Language Models (SLMs, 135M–3B parameters). We demonstrate that low-bit quantization introduces significant probability distribution divergence, reducing static speculative acceptance rates by up to 21.8%. To resolve this divergence, we propose Entropy-Gated Adaptive Speculation (EG-Spec), a dynamic lookahead policy that monitors token-level Shannon entropy and prediction margins to terminate drafting prior to catastrophic verification rejections. On rigorous empirical benchmarks spanning mathematical reasoning (GSM8K), code generation (HumanEval), and abstractive summarization, SpecEdge achieves a 1.82×–2.38× wall-clock speedup over standard auto-regressive baselines and reduces peak VRAM by 58.2%, while mathematically maintaining exact target output distributions. Non-parametric hypothesis testing (paired Wilcoxon signed-rank test, p < 0.001) validates that EG-Spec consistently outperforms static lookahead horizons on quantized edge architectures.

Zenodo (CERN European Organization for Nuclear Research)
Jain University (IN)
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

SpecEdge: Adaptive Speculative Decoding and Quantization Dynamics for Edge Small Language Models — Benedict Baah · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS