Gaurav Batule · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22995905
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
We describe DiffGQA, an attention layer that combines differential attention with grouped query attention (GQA). Each query head forms two attention distributions using separate query and key projections; their weighted difference is applied to a shared value representation. Within each group, query heads share the two key projections and one value projection. We specify the tensor shapes, normalization, causal masking, and parameter and key value (KV) cache costs. For a model with half as many KV groups as query heads and equal subhead widths, the attention projection count is 12.5% above standard multihead attention, while the KV cache stores 25% fewer scalars. These are analytical comparisons under the stated configuration, not measured speed or accuracy gains. A controlled empirical comparison remains necessary to establish whether the combination improves language modeling quality or practical efficiency.
No comments yet — start the discussion below.