SATIN: a sparse attention application-specific instruction-set processor with custom mixed-precision extensions for energy-efficient edge inference of transformer large language models

Edge inference of Transformer large language models is constrained by the quadratic cost of self-attention and by the memory traffic of the key-value cache. Existing accelerators treat sparsity and mixed precision as separate fixed-function mechanisms and provide no accuracy guarantee for an individual inference. This paper proposes a sparse attention application-specific instruction-set processor (SATIN). SATIN treats pruning as the zero-bit extreme of a bit-width continuum and thereby unifies sparsity and mixed precision as bit-plane precision. Partial sums from most-significant-bit-first serial computation supply an importance signal at low additional cost, together with a monotonically tightening error bound. One signal drives three coupled mechanisms: token sparsity on the query and key side, graded precision on the value side in proportion to the softmax weights, and importance-adaptive fetch depth for the key-value cache. A custom bit-plane instruction-set extension and a global accuracy budget allow every inference to carry a verifiable accuracy certificate. On public language and vision benchmarks the accuracy loss at an accuracy budget of one percent stays near half a percentage point. SATIN attains an effective energy efficiency of about 43.5 tera-operations per second per watt and a speedup of about 2.33 times over full precision, and reduces long-context key-value memory traffic to about one third of the dense baseline while remaining programmable.

Authors

Institutions

Publication Details

Journal
Journal of King Saud University - Computer and Information Sciences
Published
2026-09-25
DOI
https://doi.org/10.1007/s44443-026-01313-1
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

SATIN: a sparse attention application-specific instruction-set processor with custom mixed-precision extensions for energy-efficient edge inference of transformer large language models

Zhaoyu Lu
Journal of King Saud University - Computer and Information Sciences
Advanced Neural Network Applications
article

SATIN: a sparse attention application-specific instruction-set processor with custom mixed-precision extensions for energy-efficient edge inference of transformer large language models

Zhaoyu Lu
article en

Abstract

Edge inference of Transformer large language models is constrained by the quadratic cost of self-attention and by the memory traffic of the key-value cache. Existing accelerators treat sparsity and mixed precision as separate fixed-function mechanisms and provide no accuracy guarantee for an individual inference. This paper proposes a sparse attention application-specific instruction-set processor (SATIN). SATIN treats pruning as the zero-bit extreme of a bit-width continuum and thereby unifies sparsity and mixed precision as bit-plane precision. Partial sums from most-significant-bit-first serial computation supply an importance signal at low additional cost, together with a monotonically tightening error bound. One signal drives three coupled mechanisms: token sparsity on the query and key side, graded precision on the value side in proportion to the softmax weights, and importance-adaptive fetch depth for the key-value cache. A custom bit-plane instruction-set extension and a global accuracy budget allow every inference to carry a verifiable accuracy certificate. On public language and vision benchmarks the accuracy loss at an accuracy budget of one percent stays near half a percentage point. SATIN attains an effective energy efficiency of about 43.5 tera-operations per second per watt and a speedup of about 2.33 times over full precision, and reduces long-context key-value memory traffic to about one third of the dense baseline while remaining programmable.

Journal of King Saud University - Computer and Information SciencesVol. 38(8)
New York University (US)
Affordable and clean energy
Openalex Percentile: Top 14%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

SATIN: a sparse attention application-specific instruction-set processor with custom mixed-precision extensions for energy-efficient edge inference of transformer large language models — Zhaoyu Lu · Journal of King Saud University - Computer and Information Sciences (2026) | TGRS Research Map | TGRS