GNR-Q: Inference-Time State-Conditioned Reconstruction for Memory-Efficient Quantized Language Models
Low-bit quantization reduces the memory footprint of large language models (LLMs), but the accompanying loss of numerical precision perturbs internal representations and can degrade next-token prediction. We ask whether part of this lost representation fidelity can be reconstructed at inference time from the quantized model's surviving hidden state. We introduce GNR-Q (Guided Neural Reconstruction for Quantization), a compact state-conditioned sidecar trained offline to predict the residual between quantized and full-precision hidden representations while leaving the quantized backbone frozen.Using Qwen3-4B-Base, we first study reproducibility, insertion layer, and reconstructor capacity under simulated W4A16 and W3A16 quantization, then evaluate an actual TorchAO packed W4 model. Packing 252 decoder linear modules reduces measured CUDA-allocated model memory from 7.545 GiB in BF16 to 2.792 GiB, a 63.0% reduction and a 2.70x compression ratio. A 0.986M-parameter GNR-Q sidecar requires 1.88 MiB in BF16. In the primary packed experiment, reconstruction-only training reduces final hidden-state MSE from 0.5376 to 0.3901 and improves perplexity from 15.000 to 14.415, recovering 27.52% of the quantization-induced cross-entropy gap without next-token supervision. An independent matched run recovers 27.92%. In that run, a global mean-residual correction and a shuffled state-residual control recover only 16.06% and 14.72%, respectively, indicating that correct state-residual correspondence contributes substantially beyond state-independent correction. The WikiText-trained sidecar transfers without retraining to C4 validation text and recovers 35.67% of the C4 cross-entropy gap. Spectral analysis further shows that the learned correction is highly concentrated while the residual left after correction is more diffuse. In batch-1 microbenchmarks, packed W4 and packed W4 plus GNR-Q have indistinguishable prefill and decode throughput at the resolution of the measurement.The evidence supports a narrow conclusion: part of the representation error introduced by low-bit quantization is predictable from the surviving quantized state and can be reconstructed by a very small runtime module while retaining nearly all of the memory savings of packed W4 inference.
Authors
- SeungGeun Baeck
- Dumi Pyo
- HaeJung Suk
Institutions
- XLAB (Slovenia) (SI)
- Ajou University (KR)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23113127
- Primary Topic
- Machine Learning in Materials Science
- Type
- article
- Field-Weighted Citation Impact
- 0.00