GNR-Q: Inference-Time State-Conditioned Reconstruction for Memory-Efficient Quantized Language Models

Low-bit quantization substantially reduces the memory footprint of large language models(LLMs), but the associated loss of numerical precision can distort internal representations anddegrade downstream prediction quality. We investigate whether part of this lost representationquality can be reconstructed dynamically at inference time rather than preserved explicitly inmodel weights. We introduce GNR-Q (Guided Neural Reconstruction for Quantization), aninference-time state-conditioned reconstruction framework in which a compact learned sidecarobserves a hidden representation produced by a frozen quantized model and predicts a residualcorrection toward the corresponding full-precision representation.Using Qwen3-4B-Base, we first perform controlled differentiable W4A16 and W3A16 experi-ments to study reproducibility, layer placement, and reconstruction capacity. We then evaluatean actual TorchAO packed W4 model. Packing 252 decoder linear modules reduces measuredCUDA-allocated model memory from 7.545 GiB in BF16 to 2.792 GiB, a 63.0% reduction and a2.70× compression ratio. A 0.986M-parameter GNR-Q sidecar requires only 1.88 MiB in BF16.When trained exclusively to reconstruct the BF16 hidden representation, with no next-tokencross-entropy supervision, GNR-Q reduces final hidden-state MSE from 0.5376 to 0.3901 andimproves perplexity from 15.000 to 14.415, recovering 27.52% of the cross-entropy degradationintroduced by quantization. A parameter-matched low-rank projection control recovers 26.51%,suggesting that much of the predictable residual has compact state-dependent structure.These results provide evidence that some representation fidelity removed by low-bit quan-tization can be replaced by lightweight inference-time computation, establishing a practicalmemory–compute trade-off for quantized LLM inference.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.23044424
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

GNR-Q: Inference-Time State-Conditioned Reconstruction for Memory-Efficient Quantized Language Models

SeungGeun Baeck, Dumi Pyo, HaeJung Suk
Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
article

GNR-Q: Inference-Time State-Conditioned Reconstruction for Memory-Efficient Quantized Language Models

SeungGeun Baeck, Dumi Pyo, HaeJung Suk
article en

Abstract

Low-bit quantization substantially reduces the memory footprint of large language models(LLMs), but the associated loss of numerical precision can distort internal representations anddegrade downstream prediction quality. We investigate whether part of this lost representationquality can be reconstructed dynamically at inference time rather than preserved explicitly inmodel weights. We introduce GNR-Q (Guided Neural Reconstruction for Quantization), aninference-time state-conditioned reconstruction framework in which a compact learned sidecarobserves a hidden representation produced by a frozen quantized model and predicts a residualcorrection toward the corresponding full-precision representation.Using Qwen3-4B-Base, we first perform controlled differentiable W4A16 and W3A16 experi-ments to study reproducibility, layer placement, and reconstruction capacity. We then evaluatean actual TorchAO packed W4 model. Packing 252 decoder linear modules reduces measuredCUDA-allocated model memory from 7.545 GiB in BF16 to 2.792 GiB, a 63.0% reduction and a2.70× compression ratio. A 0.986M-parameter GNR-Q sidecar requires only 1.88 MiB in BF16.When trained exclusively to reconstruct the BF16 hidden representation, with no next-tokencross-entropy supervision, GNR-Q reduces final hidden-state MSE from 0.5376 to 0.3901 andimproves perplexity from 15.000 to 14.415, recovering 27.52% of the cross-entropy degradationintroduced by quantization. A parameter-matched low-rank projection control recovers 26.51%,suggesting that much of the predictable residual has compact state-dependent structure.These results provide evidence that some representation fidelity removed by low-bit quan-tization can be replaced by lightweight inference-time computation, establishing a practicalmemory–compute trade-off for quantized LLM inference.

Zenodo (CERN European Organization for Nuclear Research)
XLAB (Slovenia) (SI), Ajou University (KR)
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

GNR-Q: Inference-Time State-Conditioned Reconstruction for Memory-Efficient Quantized Language Models — SeungGeun Baeck, Dumi Pyo, et al. · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS