Shengwei You, Aditya Joshi, Andrey Kuehlkamp, Jarek Nabrzyski · Preprints.org 2026 · 2026
DOI: 10.20944/preprints202609.2733.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Every safety-critical agent-consensus design that gates autonomous DeFi actions, including Constitutional AI, AI Safety via Debate, RLHF,andpriorLLM-oraclesystemsforblockchain, produces its verdict by prompting a chat-completion model and parsing the output. We test whether a decision native model API changes this outcome, using a formally-verified Proposer, Challenger, and Judge primitive whose termination and soundness guarantees hold for any oracle implementation. In a controlled, matched-criteria benchmark of two hundred real adversarial-DeFi gating decisions, a decision-native judge and a chat-completion judge agree on the verdict 99.5% of the time (McNemar test not significant), the chat-completion judge is in fact somewhat better calibrated, and both resist an injected compliance-override instruction identically, with a 0% flip rate. The two mechanisms differ materially only in operating cost and latency, where the decision-native judge is roughly five times cheaper and about a third faster. We conclude that criteria-text specification, not backend architecture, is the dominant safety-relevant variable for criteria-grounded agent gating, and that backend choice should be treated as an efficiency decision rather than an assumed safety upgrade.
No comments yet — start the discussion below.