Noah Zelezny · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22119017
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Vector quantization (VQ) stores each group of d consecutive weights as an index into a K-entry codebook fit by k-means to the weights themselves, at log2(K)/d bits per weight, with no calibration data. We compare it against data-free round-to-nearest affine quantization (group size 64, the quantizer MLX ships) on three models: two mixture-of-experts models, Qwen3.5-397B-A17B and Qwen3.6-35B-A3B, and one dense model, Qwen3.8-27B. Only the byte-dominant tensors are quantized by the method under test — the routed experts, or the dense model's feed-forward (MLP) layers — and every other tensor is byte-identical between the two arms, so the quantizer is the only difference. Quality is exact full-vocabulary KL divergence from the bf16 model, paired over the same 12,288 positions on each of three corpora, with cluster-robust t statistics and a measured fit-to-fit noise floor for each model. Below 5 bits per weight, at matched size (the two builds within 1 GiB), VQ is never worse than affine on any corpus and is better on most. On the 397B, on the skeleton of the leading community mixed-precision build, VQ experts at that 2.6-bit build's exact size reduce divergence by 51% on prose, 42% on code and 74% on literary text, and a VQ build 11.3 GiB smaller reduces it by 31%, 11% and 44%; at 3.1 bits a VQ build 22.5 GiB smaller than the 3.5-bit build matches it. On the 35B, at identical size, VQ reduces divergence against 3-bit affine experts by 38%, 14% and 60%. On the 27B, a VQ build 0.5 GiB smaller than one with 4-bit affine MLPs is better by 18% on prose and 8% on literary text, with code within noise. From about 5.3 bits upward the two quantizers tie on both models where that range was measured, in every cell but one (35B literary text at 6.2 bits, where affine is better): each reaches the floor set by the shared skeleton. Separately, a fitter change that improves reconstruction of the largest-magnitude weights at identical size degrades literary text by 35–60% on both MoE models while improving code on one: a bulk reconstruction statistic did not rank these builds. The VQ builds measured here are published with the Metal kernels that run them. VQ is slower than affine at decode.
No comments yet — start the discussion below.