Which Interference Does 4-Bit Weight Quantization Harm? It Depends on the Operating Point
Language models are commonly served with 4-bit weights, and the format is chosen on perplexity and benchmark averages, which do not vary how strongly the context competes with the answer. A quantized model that fails to resolve such interference can answer fluently with the competing binding. The closest controlled study measured this failure with the competitor placed before the target only. Here we cross proactive and retroactive layouts with list length, interference load, and weight and KV-cache formats on recall trials paired within item in 1.5–8B open-weight models, and recompute attention in float32 on the same trials. Whether NF4 damage is layout-selective changes with the operating point within one model. On Qwen2.5-1.5B the NF4 penalty is proactive-selective on 16-pair lists (accuracy selectivity +0.231 [+0.181, +0.281]), and the format × layout interaction moves toward zero with list length (change +1.07 [+0.43, +1.71] from 16 to 64 pairs), where it is no longer resolved. On Granite-4.2-3B, at an operating point matched on full-precision accuracy alone, the layout contrast appears only under load (change −0.68 [−1.12, −0.24]). NF4 and FP4 at equal bit width leave different damage signatures, while INT8 weights cost little and 8- and 4-bit KV caches, tested at 1.5B, stay near clean. On Qwen2.5 an attention shift toward the competitor accompanies the damage and predicts which answers flip beyond the full-precision margin (held-out AUC gain +0.193 at 1.5B), yet restoring the late layers where it appears does not recover detectably more accuracy than restoring early ones. A single-layout test at one operating point therefore cannot certify a 4-bit format for interference-heavy use: both layouts have to be measured at the workload's own list length and load.
Authors
- Ya-Fen Yeh
- Guan-Yuan Chen (ORCID: https://orcid.org/0000-0003-3298-0624)
Institutions
- National Tsing Hua University (TW)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23155355
- Primary Topic
- Advanced Neural Network Applications
- Type
- preprint