Lincoln Ho, Christoforos Kachris · Applied Sciences 2026 · 2026
DOI: 10.3390/app16199626
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
As CPUs alone are often insufficiently efficient for running Vision Transformers (ViTs), contemporary computing platforms, particularly edge systems, increasingly adopt heterogeneous architectures that combine CPUs with GPUs and Neural Processing Units (NPUs). This paper presents a detailed energy assessment of individual ViT functions running on a modern hybrid processor, the AMD Ryzen AI 9 HX 370, which features a CPU, an integrated GPU (iGPU), and a dedicated Neural Processing Unit (NPU). Unlike macro-level benchmarks, our analysis takes a fine-grained approach by focusing on the most widely used ViT functions. We measure the execution time, average power draw, and idle-subtracted energy of key mathematical and structural layers such as matrix multiplications, attention mechanisms, normalization, and activations on each engine, and we profile the CPU a second time under matched INT8 quantization so that architectural effects can be separated from precision effects. Our findings reveal clear, hardware-specific trade-offs. The NPU is substantially more efficient on dense matrix multiplication (4.5× on the feed-forward projection, about 2.9× on the attention projections) and on strided convolution (up to 9.5×), and the advantage persists against the CPU at the CPU’s own cheaper precision, so it is not attributable to the NPU’s lower precision alone. The CPU remains more efficient for elementwise, activation, and normalization operators, in some cases by an order of magnitude. Two results run against common expectation: quantization was a net energy cost on fixed hardware for thirteen of sixteen isolated operators and for the fused attention block, because quantize and dequantize overhead dominates where arithmetic per element is low; and at long context, softmax consumed 26% of the attention block’s energy while performing 1.3% of its arithmetic, a divergence no operation-count model recovers. These insights provide concrete guidelines for developers designing energy-aware runtime schedulers, showing exactly when to offload specific transformer operators to maximize efficiency on hybrid client hardware.
No comments yet — start the discussion below.