SeungGeun Baeck, Dumi Pyo, HaeJung Suk · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23044424
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Low-bit quantization substantially reduces the memory footprint of large language models(LLMs), but the associated loss of numerical precision can distort internal representations anddegrade downstream prediction quality. We investigate whether part of this lost representationquality can be reconstructed dynamically at inference time rather than preserved explicitly inmodel weights. We introduce GNR-Q (Guided Neural Reconstruction for Quantization), aninference-time state-conditioned reconstruction framework in which a compact learned sidecarobserves a hidden representation produced by a frozen quantized model and predicts a residualcorrection toward the corresponding full-precision representation.Using Qwen3-4B-Base, we first perform controlled differentiable W4A16 and W3A16 experi-ments to study reproducibility, layer placement, and reconstruction capacity. We then evaluatean actual TorchAO packed W4 model. Packing 252 decoder linear modules reduces measuredCUDA-allocated model memory from 7.545 GiB in BF16 to 2.792 GiB, a 63.0% reduction and a2.70× compression ratio. A 0.986M-parameter GNR-Q sidecar requires only 1.88 MiB in BF16.When trained exclusively to reconstruct the BF16 hidden representation, with no next-tokencross-entropy supervision, GNR-Q reduces final hidden-state MSE from 0.5376 to 0.3901 andimproves perplexity from 15.000 to 14.415, recovering 27.52% of the cross-entropy degradationintroduced by quantization. A parameter-matched low-rank projection control recovers 26.51%,suggesting that much of the predictable residual has compact state-dependent structure.These results provide evidence that some representation fidelity removed by low-bit quan-tization can be replaced by lightweight inference-time computation, establishing a practicalmemory–compute trade-off for quantized LLM inference.
No comments yet — start the discussion below.