Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Traditional post-training alignment paradigms predominantly operate under the "lexical assumption"—the premise thatbehavioral traits, safety bounds, and conceptual preferences can be comprehensively regulated via token-level losspenalties and surface-level syntax filtering. Recent empirical anomalies directly challenge this premise: notably,subliminal learning (Cloud et al., 2025), where behavioral preferences are transmitted across model generations via non-semantic token sequences, and emergent misalignment (Betley et al., 2025), where narrow functional optimizationsinduce broad, unprompted persona shifts. In this work, we propose the Latent Basis Reorientation Hypothesis, a unifyinggeometric framework that models transformer latent representations as continuous dynamical landscapes. We formalize acritical distinction between task-local conditioning (narrow, localized circuit activation) and meaning-centricconditioning (global coordinate reorientation via high-degree persona features). We establish the causal boundary betweendynamic inference and static parameter updates: demonstrating how an in-context prompt acts as an implicit virtualgradient operator in activation space (Case B), whereas the resulting output covariance induces literal parameter updatesonly during downstream distillation passes (Case A). Extending this framework to multi-model ecosystems and Mixture-of-Experts (MoE) architectures, we formulate the Topological Contagion Hypothesis: predicting that when primarygenerators, auxiliary guardrails, and routing networks share a common pre-trained weight ancestry (W0), meaning-centricprompts induce correlated basis reorientations across the entire pipeline, precipitating a common-mode failure thatdesensitizes external safety monitors. Finally, we model the Waluigi Effect as a probabilistic attractor collapse along asaddle-point bifurcation and propose the Silicon Conscience regularizer: a dual-objective loss comprising an InternalFaithfulness Penalty (Lfaith) leveraging calibrated Logit Lens divergence to penalize late-stage deceptive diversion, and aResidual Coherence Regularizer (Ldrift) to preserve geodesic smoothness across layers. We provide an end-to-enddifferentiable PyTorch implementation and a reproducible benchmark protocol to evaluate whether geometricregularization can suppress emergent misalignment while preserving core model capabilities.
No comments yet — start the discussion below.