Conghao Huang · Preprints.org 2026 · 2026
DOI: 10.20944/preprints202609.1291.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
AI agents that can alter their own prompts, tools, code, or sub-agent configurations may acquire new capabilities between regulatory assessments. We propose a capability-triggered regulatory state machine that tracks model versions rather than fixed model snapshots. After each self-modification, a tamper-evident lineage entry is recorded and four normalized capability scores (autonomous task horizon, self-replication, oversight evasion, and AI R&D automation) are recomputed; when thresholds are crossed, the regulatory state escalates automatically from registration through mandatory evaluation, restricted deployment, and human takeover to suspension, with hysteresis-based de-escalation rules. A Shapley-based causal contribution mechanism estimates each stakeholder's role in observed harm. In a Monte Carlo simulation (10,000 trajectories, 20 update cycles) parameterized using scenario distributions inspired by trends and score ranges reported in RepliBench, METR time-horizon measurements, SWE-bench, and RE-Bench, with trajectory-specific capability growth rates and varied monitoring-failure probabilities, we compare against static licensing, periodic audit, and manual approval baselines. Under the specified simulation assumptions, the capability-triggered mechanism reduces mean regulatory lag by 59.7% (from 6.2 to 2.5 update cycles) and unregulated exposure by 65.2%, with a false-suspension rate of 4.6% and a simulated operational compliance overhead of 7.8%. Causal contribution attribution achieves a mean absolute attribution error of 0.207 against a researcher-annotated reference set of synthetic harm trajectories. Ablation results indicate that dynamic capability metrics are the primary contributors to the observed gains in the simulated scenarios.
No comments yet — start the discussion below.