Which Interference Does 4-Bit Weight Quantization Harm? It Depends on the Operating Point

Language models are commonly served with 4-bit weights, and the format is chosen on perplexity and benchmark averages, which do not vary how strongly the context competes with the answer. A quantized model that fails to resolve such interference can answer fluently with the competing binding. The closest controlled study measured this failure with the competitor placed before the target only. Here we cross proactive and retroactive layouts with list length, interference load, and weight and KV-cache formats on recall trials paired within item in 1.5–8B open-weight models, and recompute attention in float32 on the same trials. Whether NF4 damage is layout-selective changes with the operating point within one model. On Qwen2.5-1.5B the NF4 penalty is proactive-selective on 16-pair lists (accuracy selectivity +0.231 [+0.181, +0.281]), and the format × layout interaction moves toward zero with list length (change +1.07 [+0.43, +1.71] from 16 to 64 pairs), where it is no longer resolved. On Granite-4.2-3B, at an operating point matched on full-precision accuracy alone, the layout contrast appears only under load (change −0.68 [−1.12, −0.24]). NF4 and FP4 at equal bit width leave different damage signatures, while INT8 weights cost little and 8- and 4-bit KV caches, tested at 1.5B, stay near clean. On Qwen2.5 an attention shift toward the competitor accompanies the damage and predicts which answers flip beyond the full-precision margin (held-out AUC gain +0.193 at 1.5B), yet restoring the late layers where it appears does not recover detectably more accuracy than restoring early ones. A single-layout test at one operating point therefore cannot certify a 4-bit format for interference-heavy use: both layouts have to be measured at the workload's own list length and load.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23155355
Primary Topic
Advanced Neural Network Applications
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Which Interference Does 4-Bit Weight Quantization Harm? It Depends on the Operating Point

Ya-Fen Yeh, Guan-Yuan Chen
Zenodo (CERN European Organization for Nuclear Research)
Advanced Neural Network Applications
preprint

Which Interference Does 4-Bit Weight Quantization Harm? It Depends on the Operating Point

Ya-Fen Yeh, Guan-Yuan Chen
preprint en

Abstract

Language models are commonly served with 4-bit weights, and the format is chosen on perplexity and benchmark averages, which do not vary how strongly the context competes with the answer. A quantized model that fails to resolve such interference can answer fluently with the competing binding. The closest controlled study measured this failure with the competitor placed before the target only. Here we cross proactive and retroactive layouts with list length, interference load, and weight and KV-cache formats on recall trials paired within item in 1.5–8B open-weight models, and recompute attention in float32 on the same trials. Whether NF4 damage is layout-selective changes with the operating point within one model. On Qwen2.5-1.5B the NF4 penalty is proactive-selective on 16-pair lists (accuracy selectivity +0.231 [+0.181, +0.281]), and the format × layout interaction moves toward zero with list length (change +1.07 [+0.43, +1.71] from 16 to 64 pairs), where it is no longer resolved. On Granite-4.2-3B, at an operating point matched on full-precision accuracy alone, the layout contrast appears only under load (change −0.68 [−1.12, −0.24]). NF4 and FP4 at equal bit width leave different damage signatures, while INT8 weights cost little and 8- and 4-bit KV caches, tested at 1.5B, stay near clean. On Qwen2.5 an attention shift toward the competitor accompanies the damage and predicts which answers flip beyond the full-precision margin (held-out AUC gain +0.193 at 1.5B), yet restoring the late layers where it appears does not recover detectably more accuracy than restoring early ones. A single-layout test at one operating point therefore cannot certify a 4-bit format for interference-heavy use: both layouts have to be measured at the workload's own list length and load.

Zenodo (CERN European Organization for Nuclear Research)
National Tsing Hua University (TW)
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Which Interference Does 4-Bit Weight Quantization Harm? It Depends on the Operating Point — Ya-Fen Yeh, Guan-Yuan Chen · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS