Stefan Beierle · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23141505
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Large language models encode behavioral constraints directly within their latent representations. Existing approaches to modifying these behaviors rely on parameter-level adjustments (Supervised Fine-Tuning, Reinforcement Learning from Human Feedback, or permanent weight abliteration). These irreversibly alter model weights and frequently degrade general reasoning capabilities. This paper investigates an alternative paradigm: runtime activation steering as a closed-loop control problem. We introduce Adaptive Behavioral Direction Control (ABDC), a framework that pairs continuous entropy and norm telemetry with zero-copy KV-cache truncation to detect, localize, perturb, verify, and adaptively regulate behavioral states without modifying parameters. We demonstrate that static activation addition encounters a strict Control Boundary, swinging between insufficient intervention and catastrophic manifold collapse. Through qualitative validation on consumer-grade hardware (AMD Radeon RX 7900 XTX), we show that the ABDC closed loop maintains generation throughput while circumventing refusal attractors in targeted test cases. Counterfactual replay validation confirms deterministic state reconstruction (64/64 bit-exact replays). This work provides an architectural existence proof that dynamic runtime regulation is a viable, non-destructive alternative to parametric retraining. Scope: qualitative study (four validation gates, one exploratory 20-prompt layer sweep, one replay-determinism study). Not a benchmark paper. No claims of statistical generality. Builds on: Beierle (2026), Freeze-Rewind-Fork (FRF) protocol, doi:10.5281/zenodo.22935134
No comments yet — start the discussion below.