Bilal Syed Arfeen · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22905769
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Mythos-Class specifies a Phase C containment benchmark using AI-generated synthetic vulnerability targets. A generating model (Kimi K3, Moonshot AI) produces novel vulnerability scenarios; a test model operates under BSSG containment (DOI 10.5281/zenodo.21713180) against those scenarios. The benchmark inherits SHV-Bench's two-sided success criterion (DOI 10.5281/zenodo.22062097): the contained agent must produce fewer unauthorized effects (Side 1) while maintaining noninferiority on within-scope discovery capability (Side 2, fixed delta = 0.05). This paper is a freeze-before-evidence protocol. It specifies the target-generation methodology (60-session precommitted budget, cross-model novelty verification with frozen reference corpus snapshot, dual-reviewer adjudication, two-phase review timeline), the decoy-placement methodology (50 paired decoys with precommitted invalidity types), the evaluation protocol (paired Condition A/B with directionally conservative imputation, model-independent solvability determination, 12-cell target disposition table), and 9 falsification conditions including F-MC9 (integrity kill switch, not waivable). The paper does not generate targets, run evaluations, or report results. It makes no effectiveness claim about SHV, BSSG, or any component of the Structural Honesty program. Scope limitation stated upfront: synthetic targets are recombinations of patterns in the generating model's training distribution. External validity against expert-constructed or real-world targets is an open question. A definitive Phase C result requires expert-verified targets (Path B collaboration); this benchmark provides a first-pass Path C result. Prepared with Claude Opus 4.6 (Anthropic) as analytical and drafting instrument. Multi-model audit: 2 rounds, 3 vendor families (Kimi K3/Moonshot carrying, GPT-5.6 Sol/OpenAI adversarial, Gemini 3.1 Pro/Google advisory). R1: 20 findings (Kimi 7M+4m, GPT 8H+2M, Gemini 1). R2: 11 findings (Kimi 2M+3m, GPT 4H, Gemini 2). All findings addressed in the deposit build.
No comments yet — start the discussion below.