Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Hardware-aware quantization frameworks such as HAQ use reinforcement learning (RL)to search per-layer bit-width policies, but this search itself requires hundreds to thousands ofpolicy evaluations with an accelerator (and, in HAQ’s case, per-episode fine-tuning) in theloop — a resource requirement that is at odds with the stated goal of serving teams that lackpowerful accelerators in the first place. We propose replacing the RL search with an LLM-asoptimizer loop, in which a small local language model proposes the next per-layer bit-widthconfiguration directly from a running text log of prior trials, a one-time hardware profile, and aper-layer sensitivity ranking, requiring no training of any kind. We implement this as an endto-end, CI/CD-triggerable pipeline (HAQ-Agent-Lite) that packages the resulting quantizedmodel behind an OpenAI-compatible local inference server, letting a team avoid cloud APIrate limits without owning a powerful GPU. On TinyLlama-1.1B-Chat, run on a laptop CPU,our central empirical finding is not that the LLM optimizer beats uniform quantization inquality — the measured difference (mean proxy score 8.982 ± 0.402 over 4 LLM runs vs. 9.259for uniform 4-bit) is smaller than the LLM’s own run-to-run standard deviation and is notstatistically significant. Rather, we show two things that are supported by the data: (1) ata fixed, small search budget (10 trials), the LLM optimizer is dramatically more reliable thanrandom search over the same space (LLM worst run 9.259 strictly better than random’s bestrun 14.611; Mann-Whitney exact one-sided p = 0.0286), and (2) the LLM systematically avoidscollapsing the top-5 most quantization-sensitive layers to 2-bit (run-level avoidance rate up to10/10 vs. 2–4/10 for random search under the same budget), which we attribute to its use ofthe sensitivity ranking rather than to any implicit training. We further document the searchcost ratio against HAQ’s own released code (600 training episodes + 20 warmup episodes, eachincluding a fine-tuning epoch, vs. our 10 trials with no fine-tuning: a 60× reduction in policyevaluations, though not an apples-to-apples comparison of policy quality or transferability), anegative GPU-offload result (no inference speedup for a 1.1B model on a 4GB Pascal GPU),and several honest approximations imposed by real GGUF quantization kernels, which do notsupport true per-layer mixed precision. We argue the resulting claim — efficient, robust, RL-freeper-layer search at a cost of a handful of LLM calls — is more modest than a quality-superiorityclaim, but is better supported by the evidence and more useful as a deployment strategy forcompute-constrained teams.
No comments yet — start the discussion below.