Sri Pranav Kunda, Alexander Kurz, Tomáš Dominik, Uri Maoz · Figshare 2026 · 2026
DOI: 10.6084/m9.figshare.33782812.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Existing activation steering methods exploit interpretable directions in the residual stream to manipulate model behavior, while finetuning methods offer greater behavioral control at the cost of interpretability. Prior work has shown that activation steering in large language models can be simulated by modifying their weights. However, the converse direction, representing the more expressive weight adaptations as activation steering, remains less understood. For this, we introduce \textsc{first-order steering}, a formulation of activation steering as a first-order approximation of each layer function with respect to a strength vector $\vs$. When considering a family of models indexed by a steering strength vector $\vs$, each component $s_i$ naturally induces a composable direction in activation space, which we use to compute activation steering vectors. We also establish theoretical bounds on the approximation error of first-order steering, through which we introduce \textsc{HeRD-Merging}, a model-merging procedure for minimizing these bounds while maintaining continuous and composable behavioral control. Thus, first-order steering offers a quantifiable indicator of steering performance which existing steering methods lack. Empirically, steering vectors derived from our formulation are more consistently accurate than existing steering methods. Furthermore, HeRD-Merging more faithfully translates weight-space adaptation into activation-space interventions while matching the capabilities of conventional merging baselines.
No comments yet — start the discussion below.