Benoît Dherin, Michael Munn, Xavi Gonzalvo, Adrian Goldwaser, Blaž Bratanič, Ananth Balashankar, Andrey Vlasov, Pinzhi Huang, Cécile Logé, Nicole Mitchell, André Fernandes, Trilok Acharya, Wendy Kan, Ziyue Wang, Hanna Mazzawi, Felipe Tiengo Ferreira, Mor Geva · arXiv (Cornell University) 2026 · 2026
DOI: 10.48550/arxiv.2609.36434
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Context tokens in a transformer-based language model can be absorbed into the model's weights as a multiplicative operator. We study this operator in the setting of safety instructions and show that influencing its dominant eigenvalue modulates how strongly the instruction shapes generation. We derive a Contrastive Safety Loss with a suppression weight that controls the tradeoff between emphasizing the safety instruction on harmful queries while suppressing it on harmless queries. Varying the suppression weight maps a relationship between the attack success and the over-refusal rates, supporting the hypothesis that the operator's eigenvalue acts as a continuous dial for the instruction's influence. Moreover, this relationship holds relatively independently of how the Safety Loss is parameterized, yielding Pareto-improved safety instructions for appropriate values of suppression weight.
No comments yet — start the discussion below.