MARY VARSHITHA, Mohammed Ayaan · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22978297
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
PIVOT (Prospective Causal Attribution for AI Agents) is a controlled benchmark for evaluating whether an AI system can predict the causal impact of a pending action before that action is executed. The benchmark follows a prospective decision-making loop: observe the current state, identify a pending action, predict its causal effect, execute the action, replay the same state under an alternative action, and compare the resulting outcomes. This separates prospective causal prediction from purely retrospective attribution. PIVOT evaluates models across controlled scenarios designed to test different causal structures, including immediate effects, delayed effects, hidden pivots, recoverable outcomes, redundant actions, false causal cues, and interaction effects. The framework uses intervention-based replay to obtain counterfactual ground truth while keeping the evaluation setting controlled and reproducible. The accompanying report documents the benchmark design, experimental methodology, available results, limitations, audit corrections, and reproducibility considerations. Some previously reported metrics could not be independently verified from the currently available evidence and are therefore explicitly identified as unverifiable rather than reconstructed or inferred. This work is intended as a research benchmark and methodological study for prospective causal attribution in AI agents, rather than as a claim of general causal reasoning ability in real-world AI systems.
No comments yet — start the discussion below.