Hazama Kaizuka · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23020503
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This technical note reports an empirical comparison of CPU and Vulkan GPU inference for quantized GGUF large language models on a consumer Android edge device. Measurements were performed on a OnePlus Nord 3 equipped with a MediaTek Dimensity 9000 SoC and Mali-G710 MC10 GPU, using a fixed llama.cpp revision and a Vulkan 1.1-compatible execution path. The study compares multiple model scales and architectures, including LFM2.5, Ministral 3, Granite 4.0 Micro, Qwen3.5, and the LFM2.5-8B-A1B Mixture-of-Experts model. The results show that mobile GPU inference does not automatically outperform ARM CPU inference. CPU/GPU performance varies significantly with model architecture and inference phase, particularly between prefill and autoregressive generation. The measurements also suggest that total parameter count alone is insufficient to predict CPU/GPU crossover behavior on mobile edge hardware. The note discusses implications for edge AI and physical AI systems, where heterogeneous CPU/GPU resource allocation may be more useful than unconditional preference for GPU execution.
No comments yet — start the discussion below.