Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Modern AI models, from language models to predictive encoders, learn associations from their training data. Most evaluations of their reasoning check whether an answer matches a set pattern rather than whether it follows how a system really works. For systems that change over time, that matters. An answer can be right while the dynamics behind it are impossible. This paper describes an evaluated integration of established methods that addresses this problem, drawing on Pearl’s distinction between observing and intervening, equation discovery, the stock-and-flow representation of feedback and delay, LLM-as-judge evaluation, and the data engineering needed to capture interventions. The result is a Causal Generator: an executable model of hypothesized causal dynamics, recovered from interventional data and validated within a declared intervention domain, against which specified claims about a system’s behavior can be checked. On a benchmark ranging from a single tank to chaotic dynamics, frontier models were far better at verifying claims than at forecasting, and the tested frontier judges given the generator’s predicted trajectory matched or exceeded their other configurations at every level. Generators need interventional data and real understanding of the mechanism, and they have their own failure modes. But, where they fit, they complement current answer-level and judge-based evaluation techniques, and can be a strong addition to a larger evaluation framework.
No comments yet — start the discussion below.