Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
I run Qwen3.8-27B, in Unsloth's UD-IQ4_XS quantisation (14.25 GB, about 4.2 bits per weight), entirely on a 16 GB AMD RX 7800 XT with a context of 104,448 tokens. It decodes at 31 tokens/s on short prompts and 22 tokens/s with the context nearly full. This works mostly because the model keeps a KV cache on only 16 of its 64 layers, but the last few thousand tokens required measuring VRAM use down to the mebibyte. I describe that budget, a formula that predicts the peak to within 1 MiB, and a failure that appears about 80 MiB before the card is full. The paper also covers asymmetric KV-cache quantisation, micro-batch size, Vulkan vs. ROCm, speculative decoding and a CPU-side vision encoder, and includes the full llama-server command.
No comments yet — start the discussion below.