Component-aware self-speculative decoding for hybrid language models: An architectural viability study

Speculative decoding reduces language-model inference latency when a cheap draft distribution matches the target. Existing self-speculative methods skip layers in homogeneous Transformer stacks; we ask whether hybrid language models can instead draft from their internal state-space or linear-attention components as parameter-free drafters. We introduce component-aware self-speculative decoding and evaluate it on nine models: eight hybrids across five families and both integration paradigms, plus a pure-Transformer control. Under greedy decoding the parallel hybrid Falcon-H1-0.5B reaches a mean acceptance rate A 2 = 0.680 , whereas the sequential hybrid Qwen3.5-0.8B reaches only 0.038. The governing factor is not the parallel/sequential label but functional attention-dependence : acceptance falls as the perplexity increase under attention suppression grows (Spearman ρ = − 0.79 , p = 0.028 over the eight hybrids; ρ = − 0.75 with the control), and an attention-light sequential hybrid (granite) outperforms a parallel one (Falcon-H1-1.5B). The separation is stable at the 0.5B and 3B scales tested (permutation test, p = 0.962 ). Acceptance viability is thus demonstrated, but it does not by itself yield speedup. Using hardware-measured draft/verify cost ratios ( c eff ), and even after activating the official accelerated state-space kernels, no drafter reaches a theoretical speedup above one (best S theory = 0.986 ). With the kernels active the binding constraint is acceptance rather than drafting cost, which has fallen to its parameter-count floor. We frame the work as an architectural viability study: the internal state-space drafter is acceptance-viable in the right architectures, but turning it into wall-clock speedup is a separate execution-stack and modeling problem.

Authors

Institutions

Publication Details

Journal
Computers & Electrical Engineering
Published
2026-09-14
DOI
https://doi.org/10.1016/j.compeleceng.2026.111505
Primary Topic
Generative Adversarial Networks and Image Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Component-aware self-speculative decoding for hybrid language models: An architectural viability study

Guillermina Tormo‐Carbó, Elies Seguí‐Mas, Hector Borobia
Computers & Electrical Engineering
Generative Adversarial Networks and Image Synthesis
article

Component-aware self-speculative decoding for hybrid language models: An architectural viability study

Guillermina Tormo‐Carbó, Elies Seguí‐Mas, Hector Borobia
article en

Abstract

Speculative decoding reduces language-model inference latency when a cheap draft distribution matches the target. Existing self-speculative methods skip layers in homogeneous Transformer stacks; we ask whether hybrid language models can instead draft from their internal state-space or linear-attention components as parameter-free drafters. We introduce component-aware self-speculative decoding and evaluate it on nine models: eight hybrids across five families and both integration paradigms, plus a pure-Transformer control. Under greedy decoding the parallel hybrid Falcon-H1-0.5B reaches a mean acceptance rate A 2 = 0.680 , whereas the sequential hybrid Qwen3.5-0.8B reaches only 0.038. The governing factor is not the parallel/sequential label but functional attention-dependence : acceptance falls as the perplexity increase under attention suppression grows (Spearman ρ = − 0.79 , p = 0.028 over the eight hybrids; ρ = − 0.75 with the control), and an attention-light sequential hybrid (granite) outperforms a parallel one (Falcon-H1-1.5B). The separation is stable at the 0.5B and 3B scales tested (permutation test, p = 0.962 ). Acceptance viability is thus demonstrated, but it does not by itself yield speedup. Using hardware-measured draft/verify cost ratios ( c eff ), and even after activating the official accelerated state-space kernels, no drafter reaches a theoretical speedup above one (best S theory = 0.986 ). With the kernels active the binding constraint is acceptance rather than drafting cost, which has fallen to its parameter-count floor. We frame the work as an architectural viability study: the internal state-space drafter is acceptance-viable in the right architectures, but turning it into wall-clock speedup is a separate execution-stack and modeling problem.

Computers & Electrical EngineeringVol. 140
Artificial Intelligence Research Institute (ES), Universitat Politècnica de València (ES)
Quality Education
Openalex Percentile: Top 13%
Generative Adversarial Networks and Image Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.