Mixed-Precision Quantization for Language Models: Techniques and Prospects

The rapid scaling of Language Models (LMs) has resulted in unprecedented computational, memory, and energy requirements, making their training and deployment increasingly unsustainable. Quantization has emerged as a crucial compression technique for reducing model size, alleviating memory bottlenecks, and accelerating inference. However, while uniform low-bit quantization (e.g., INT8, INT4) provides significant efficiency gains, it can degrade accuracy in sensitive components of transformer-based LMs. Mixed-precision quantization offers a promising alternative by selectively allocating precision across layers or within tensors to strike a balance between efficiency and accuracy. This survey provides a comprehensive overview of Mixed-Precision quantization frameworks for LMs (MXPLMs). We first review quantization fundamentals, including uniform and non-uniform quantizers, quantization granularity, and methods widely used in post-training quantization. We then categorize and compare recent MXPLM frameworks according to their bit allocation strategies and precision configurations across weights, activations, and key-value caches. A comparative analysis highlights differences in perplexity, zero-shot task performance, and deployment trade-offs. Furthermore, we contrast MXPLMs with earlier mixed-precision quantization methods for deep neural networks, identifying strategies that transfer and those that face challenges in the LM setting. Then, we discuss quantization-compatible hardware, KV cache and mixture-of-experts quantization, and summarize open issues and future directions, including hardware-aware design, activation quantization, and scalable optimization methods for billion-parameter models.

Authors

Institutions

Publication Details

Journal
ACM Computing Surveys
Published
2026-09-16
DOI
https://doi.org/10.1145/3848510
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Mixed-Precision Quantization for Language Models: Techniques and Prospects

Olga Krestinskaya, Fadi Kurdahi, Ahmed M. Eltawil, Marios Fournarakis et al.
ACM Computing Surveys
Advanced Neural Network Applications
article

Mixed-Precision Quantization for Language Models: Techniques and Prospects

Olga Krestinskaya, Fadi Kurdahi, Ahmed M. Eltawil, Marios Fournarakis, Jinane Bazzi, K. Saláma, Mariam Rakka, Mohammed E. Fouda
article en

Abstract

The rapid scaling of Language Models (LMs) has resulted in unprecedented computational, memory, and energy requirements, making their training and deployment increasingly unsustainable. Quantization has emerged as a crucial compression technique for reducing model size, alleviating memory bottlenecks, and accelerating inference. However, while uniform low-bit quantization (e.g., INT8, INT4) provides significant efficiency gains, it can degrade accuracy in sensitive components of transformer-based LMs. Mixed-precision quantization offers a promising alternative by selectively allocating precision across layers or within tensors to strike a balance between efficiency and accuracy. This survey provides a comprehensive overview of Mixed-Precision quantization frameworks for LMs (MXPLMs). We first review quantization fundamentals, including uniform and non-uniform quantizers, quantization granularity, and methods widely used in post-training quantization. We then categorize and compare recent MXPLM frameworks according to their bit allocation strategies and precision configurations across weights, activations, and key-value caches. A comparative analysis highlights differences in perplexity, zero-shot task performance, and deployment trade-offs. Furthermore, we contrast MXPLMs with earlier mixed-precision quantization methods for deep neural networks, identifying strategies that transfer and those that face challenges in the LM setting. Then, we discuss quantization-compatible hardware, KV cache and mixture-of-experts quantization, and summarize open issues and future directions, including hardware-aware design, activation quantization, and scalable optimization methods for billion-parameter models.

ACM Computing Surveys
Al Ain University (AE), University of California, Irvine (US), Kootenay Association for Science & Technology (CA), Irvine University (US), King Abdullah University of Science and Technology (SA)
Openalex Percentile: Top 99%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.