Akan Mukhametgali · OSF Preprints (OSF Preprints) 2026 · 2026
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This research investigates the geometric structure of safety alignment mechanisms within Vision-Language Models, specifically examining whether refusal behaviors triggered by visual inputs are mediated by the same low-dimensional subspace as those triggered by textual inputs. The study formalizes refusal in a VLM as a union of modality-conditioned cones, comprising a text-based refusal cone and a visual refusal cone. By analyzing the principal angles between these subspaces across multiple layers and model scales, the research demonstrates that the visual refusal cone consistently retains a significant energy component that is orthogonal to the text refusal cone. This geometric separation is shown to be scale-robust, meaning it does not diminish as the model size increases, and is particularly pronounced in mixture-of-experts architectures. To establish a causal understanding of this phenomenon, the study employs a controlled experimental setup wherein safety alignment is injected into a base, non-aligned VLM using a language-only Low-Rank Adaptation. The findings reveal a critical transfer gap: while the model achieves complete refusal on harmful textual prompts, its refusal rate for harmful visual inputs drops significantly, exposing a behavioral vulnerability. This indicates that safety supervision restricted to the text modality fails to fully project the refusal mechanism onto the visual processing pathways. The research further explores the implications of this vulnerability for cross-modal abliteration attacks. Although the representational separation suggests that text-only defenses leave a structural gap exploitable by attackers, the empirical evaluation of cross-modal abliteration at the tested scale did not yield a statistically robust functional attack advantage, presenting an honest negative result regarding immediate attack efficacy. Ultimately, the study proposes and validates a constructive defense mechanism. By incorporating modality-complete supervision that explicitly aligns both the text and visual refusal cones, the behavioral gap is successfully closed, resulting in comprehensive refusal capabilities across both modalities. The central conclusion is that safety alignment and abliteration defenses engineered exclusively within the text representation inherently inherit a persistent visual vulnerability. To ensure robust security in multimodal environments, alignment protocols must explicitly address and harmonize the distinct geometric subspaces governing visual and textual refusal mechanisms.
No comments yet — start the discussion below.