Stefano Fazzino · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23166741
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This work investigates how post-training weight quantization affects the knowledge memorized by a language model. To this end, we train a transformer with 1.1 million parameters on synthetic biographies, a setting in which the amount of information is known exactly, and we vary both the information load and the number $K$ of values each attribute can take. After training, the weights are rounded to $b$ bits and the surviving knowledge is measured fact by fact. We show that the number of bits needed to retain 99\% of the facts grows from 4 to 8 with the load, following the logarithmic dependence on the distance from capacity that Gardner's theory predicts for the perceptron. The narrowing of the margins is accompanied by a growth of the perturbation, which a linear propagation of the error through the network attributes mainly to the increase of the gain with which the subsequent layers carry the rounding error to the output, rather than to a coarser quantization grid. The fraction of facts lost at each precision is predicted, with no parameter fitted to the losses, by a survival law in which the correct answer must stay above all its competitors under a perturbation partly shared among them, whereas a description based on the runner-up alone underestimates the losses, increasingly so as $K$ grows. We derive the shared component from the geometry of the output layer and build a solvable model of a multi-class readout, which agrees with simulations without free parameters and reproduces the trends observed in the transformer. On three pretrained language models the structure of the perturbation is confirmed, but the losses are overestimated: the facts these models know are held with small margins, and the perturbation has heavy tails and affects facts to different extents depending on the directions from which they are read, two effects that together reduce the discrepancy to less than half.
No comments yet — start the discussion below.