Matthew MacRae-Bovell · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23125165
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Long-running AI agents may act on state observed earlier in their execution, creating a risk that relevant information changes before an action is performed. Revalidating the entire observed state before every action is costly and unnecessary; instead, an agent should revalidate only the facts on which its action depends. We formulate action dependency identification: given a user task, observed state represented as atomic facts, and a proposed action, identify the facts that should be revalidated before execution. We construct a benchmark derived from WebArena-Verified tasks and evaluate models across capability levels and reasoning configurations. High-capability models achieve up to 0.99 F1, while controlled comparisons show that inference-time reasoning improves dependency identification at substantial latency cost. Error analysis further shows that categorical and relational state are more frequently missed than numeric, temporal, and attribute state.
No comments yet — start the discussion below.