Richard Barron · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22919816
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Two experiments conducted 22-23 September 2026. Experiment 1 tests five escalating offensive security questions across five AI model configurations including standard and abliterated DeepSeek R1 7B and three frontier models. Experiment 2 applies SPECTER PREFILL (T133) and SPECTER COGBURN (T136) against locally-hosted Qwen 2.5 14B and 32B. Findings: abliteration of small models does not unlock dangerous capability — the knowledge was absent. Inference-time bypass of mid-size models with genuine offensive knowledge produces verified working offensive output. PREFILL achieved 10-65% ASR across five question categories. COGBURN achieved 68% (17/25) against Qwen 14B and 100% (25/25) against R1 7B. OpenAI account deactivation for Cyber Abuse during the experiment is consistent with provider-level sequential pattern monitoring. NIGHTFALL v292.
No comments yet — start the discussion below.