Abena Primo, Azubike Okpalaeze, Al Amin, Rohan Thompson · · 2026
DOI: 10.68414/bdbh8187
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Frontier artificial intelligence laboratories now publish detailed accounts of how their models behave under adversarial pressure: red-team transcripts, capability thresholds, interpretability findings, and responsible-scaling frameworks. These disclosures are widely treated as safety infrastructure, built to keep dangerous capabilities contained and visible to regulators. This paper argues that the same disclosures can function as a second, less examined kind of infrastructure: a map of where a model’s offensive ceiling sits, available to anyone willing to read it. Using a comparative case study of the published capability-disclosure frameworks of Anthropic, OpenAI, and Google DeepMind, three labs operating under a shared 2024 Seoul Summit safety mandate, this paper develops a concept called the Safety Exposure Paradox, in which the transparency that makes frontier AI governable also narrows the search space for adversaries probing model behavior. The paper then traces this dynamic into financial infrastructure, where fraud detection, anti-money-laundering screening, identity verification, and algorithmic trading increasingly depend on AI systems whose upstream governance, training data, and safety ceilings sit outside any single institution’s control. Drawing on financial-stability research, dual-use disclosure literature, and recent reporting on AI-enabled synthetic identity fraud, the paper proposes a tiered disclosure model with concrete implementation steps, and a set of policy recommendations for regulators and fintech institutions managing inherited AI risk. The aim is not to argue against safety disclosure, but to show that disclosure decisions carry a financial-sector externality that current governance frameworks do not yet price in.
No comments yet — start the discussion below.