SATLLM: A Sparsity-aware Accelerator for Ternary-Weight Large Language Models

Large language models (LLMs) exhibit strong performance across applications, but their inference is computationally intensive, posing significant challenges for edge deployment. Quantization is among the most effective and widely used optimizations. In particular, ternary-weight quantization further lowers compute cost by replacing many multiplications with lightweight additions, thereby reducing complexity and energy. However, existing GPU/ASIC solutions often exhibit low utilization for ternary-weight LLMs and fail to exploit their inherent sparsity, leading to inefficient execution. To address this, we propose SATLLM, a lightweight accelerator optimized for ternary weights and sparsity. At the algorithm level, we propose a co-optimization method based on unstructured pruning and ternary-weight clustering to effectively enhance the sparsity. To further exploit this sparsity, we design a hardware-friendly sparsity-aware merging scheme and develop a customized BitLinear engine to support efficient sparse computation. In addition, we optimize the dataflow between the BitLinear engine and the GEMM unit, effectively alleviating performance bottlenecks caused by workload imbalance. Together, these optimizations significantly improve the throughput and energy efficiency of ternary-weight LLM inference. The experiments show that our proposed accelerator achieves an energy efficiency of 9.42 TOPS/W, with improvements of up to 42.41 ×, 16.99 ×, 4.42 ×, 2.49 ×, and 1.90 × compared to NVIDIA V100, A100, ANT, MECLA, and RMFA, respectively.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Design Automation of Electronic Systems
Published
2026-09-24
DOI
https://doi.org/10.1145/3844511
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

SATLLM: A Sparsity-aware Accelerator for Ternary-Weight Large Language Models

Li Shang, Changxu Liu, Fan Yang, Yifan Song et al.
ACM Transactions on Design Automation of Electronic Systems
Topic Modeling
article

SATLLM: A Sparsity-aware Accelerator for Ternary-Weight Large Language Models

Li Shang, Changxu Liu, Fan Yang, Yifan Song, Mengyi Chen, Yifeng Yang, Hengjie Cao
article en

Abstract

Large language models (LLMs) exhibit strong performance across applications, but their inference is computationally intensive, posing significant challenges for edge deployment. Quantization is among the most effective and widely used optimizations. In particular, ternary-weight quantization further lowers compute cost by replacing many multiplications with lightweight additions, thereby reducing complexity and energy. However, existing GPU/ASIC solutions often exhibit low utilization for ternary-weight LLMs and fail to exploit their inherent sparsity, leading to inefficient execution. To address this, we propose SATLLM, a lightweight accelerator optimized for ternary weights and sparsity. At the algorithm level, we propose a co-optimization method based on unstructured pruning and ternary-weight clustering to effectively enhance the sparsity. To further exploit this sparsity, we design a hardware-friendly sparsity-aware merging scheme and develop a customized BitLinear engine to support efficient sparse computation. In addition, we optimize the dataflow between the BitLinear engine and the GEMM unit, effectively alleviating performance bottlenecks caused by workload imbalance. Together, these optimizations significantly improve the throughput and energy efficiency of ternary-weight LLM inference. The experiments show that our proposed accelerator achieves an energy efficiency of 9.42 TOPS/W, with improvements of up to 42.41 ×, 16.99 ×, 4.42 ×, 2.49 ×, and 1.90 × compared to NVIDIA V100, A100, ANT, MECLA, and RMFA, respectively.

ACM Transactions on Design Automation of Electronic Systems
Fudan University (CN), Shanghai Fudan Microelectronics (China) (CN)
Affordable and clean energy
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

SATLLM: A Sparsity-aware Accelerator for Ternary-Weight Large Language Models — Li Shang, Changxu Liu, et al. · ACM Transactions on Design Automation of Electronic Systems (2026) | TGRS Research Map | TGRS