GNR-Q: Inference-Time State-Conditioned Reconstruction for Memory-Efficient Quantized Language Models

Low-bit quantization reduces the memory footprint of large language models (LLMs), but the accompanying loss of numerical precision perturbs internal representations and can degrade next-token prediction. We ask whether part of this lost representation fidelity can be reconstructed at inference time from the quantized model's surviving hidden state. We introduce GNR-Q (Guided Neural Reconstruction for Quantization), a compact state-conditioned sidecar trained offline to predict the residual between quantized and full-precision hidden representations while leaving the quantized backbone frozen.Using Qwen3-4B-Base, we first study reproducibility, insertion layer, and reconstructor capacity under simulated W4A16 and W3A16 quantization, then evaluate an actual TorchAO packed W4 model. Packing 252 decoder linear modules reduces measured CUDA-allocated model memory from 7.545 GiB in BF16 to 2.792 GiB, a 63.0% reduction and a 2.70x compression ratio. A 0.986M-parameter GNR-Q sidecar requires 1.88 MiB in BF16. In the primary packed experiment, reconstruction-only training reduces final hidden-state MSE from 0.5376 to 0.3901 and improves perplexity from 15.000 to 14.415, recovering 27.52% of the quantization-induced cross-entropy gap without next-token supervision. An independent matched run recovers 27.92%. In that run, a global mean-residual correction and a shuffled state-residual control recover only 16.06% and 14.72%, respectively, indicating that correct state-residual correspondence contributes substantially beyond state-independent correction. The WikiText-trained sidecar transfers without retraining to C4 validation text and recovers 35.67% of the C4 cross-entropy gap. Spectral analysis further shows that the learned correction is highly concentrated while the residual left after correction is more diffuse. In batch-1 microbenchmarks, packed W4 and packed W4 plus GNR-Q have indistinguishable prefill and decode throughput at the resolution of the measurement.The evidence supports a narrow conclusion: part of the representation error introduced by low-bit quantization is predictable from the surviving quantized state and can be reconstructed by a very small runtime module while retaining nearly all of the memory savings of packed W4 inference.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23113127
Primary Topic
Machine Learning in Materials Science
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

GNR-Q: Inference-Time State-Conditioned Reconstruction for Memory-Efficient Quantized Language Models

SeungGeun Baeck, Dumi Pyo, HaeJung Suk
Zenodo (CERN European Organization for Nuclear Research)
Machine Learning in Materials Science
article

GNR-Q: Inference-Time State-Conditioned Reconstruction for Memory-Efficient Quantized Language Models

SeungGeun Baeck, Dumi Pyo, HaeJung Suk
article en

Abstract

Low-bit quantization reduces the memory footprint of large language models (LLMs), but the accompanying loss of numerical precision perturbs internal representations and can degrade next-token prediction. We ask whether part of this lost representation fidelity can be reconstructed at inference time from the quantized model's surviving hidden state. We introduce GNR-Q (Guided Neural Reconstruction for Quantization), a compact state-conditioned sidecar trained offline to predict the residual between quantized and full-precision hidden representations while leaving the quantized backbone frozen.Using Qwen3-4B-Base, we first study reproducibility, insertion layer, and reconstructor capacity under simulated W4A16 and W3A16 quantization, then evaluate an actual TorchAO packed W4 model. Packing 252 decoder linear modules reduces measured CUDA-allocated model memory from 7.545 GiB in BF16 to 2.792 GiB, a 63.0% reduction and a 2.70x compression ratio. A 0.986M-parameter GNR-Q sidecar requires 1.88 MiB in BF16. In the primary packed experiment, reconstruction-only training reduces final hidden-state MSE from 0.5376 to 0.3901 and improves perplexity from 15.000 to 14.415, recovering 27.52% of the quantization-induced cross-entropy gap without next-token supervision. An independent matched run recovers 27.92%. In that run, a global mean-residual correction and a shuffled state-residual control recover only 16.06% and 14.72%, respectively, indicating that correct state-residual correspondence contributes substantially beyond state-independent correction. The WikiText-trained sidecar transfers without retraining to C4 validation text and recovers 35.67% of the C4 cross-entropy gap. Spectral analysis further shows that the learned correction is highly concentrated while the residual left after correction is more diffuse. In batch-1 microbenchmarks, packed W4 and packed W4 plus GNR-Q have indistinguishable prefill and decode throughput at the resolution of the measurement.The evidence supports a narrow conclusion: part of the representation error introduced by low-bit quantization is predictable from the surviving quantized state and can be reconstructed by a very small runtime module while retaining nearly all of the memory savings of packed W4 inference.

Zenodo (CERN European Organization for Nuclear Research)
XLAB (Slovenia) (SI), Ajou University (KR)
Openalex Percentile: Top 26%
Machine Learning in Materials Science
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.