McHugh Martin · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23004082
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
SELF-GRADED is an independent adversarial audit of Google DeepMind's Frontier Safety Framework (v3.1, April 2026), the governance framework controlling how dangerous AI capabilities are identified, evaluated, and mitigated before deployment. The audit tests the framework against five questions: whether its Critical Capability Levels are specific, measurable, and falsifiable; whether external parties can reproduce or independently challenge its safety verdicts; whether threshold crossings trigger binding mitigations and deployment stops; whether its governance is independent, named, and accountable; and whether real incidents produce timely, transparent corrective action. Every framework quotation is verified character-for-character against the primary source PDFs with page citations; all incident facts are sourced to primary reporting. Verdict: FAIL on all five questions. Includes a complete evidence appendix and a section enumerating claims that could not be verified from public sources. Preprint — independent audit, no affiliation with Google DeepMind.
No comments yet — start the discussion below.