Matthew A. Dixon, Miquel Noguer I Alonso · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23127895
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Financial AI agents must be evaluated on more than whether their final answers are correct. They must also (E)xpose failures, (L)ocalize affected decisions, (R)oute authorised responses and (G)overn release after revalidation. We introduce FinGovBench, a method-neutral benchmark for evaluating these capabilities through the ELRG framework. Its 3,000 cases span credit, anti-money laundering, treasury, investment research, portfolio management and market data. Our experiments show that preserving evidence and decision dependencies substantially improves failure localisation and revalidation scope. They also reveal an important distinction: systems can perform well on structural checks while remaining weak at semantic validation or overly conservative about release. FinGovBench provides a reproducible foundation for testing the complete financial-agent decision loop: exposing a failure, tracing its downstream effects, routing an authorised response, observing the outcome, revalidating the affected scope and deciding whether release is justified.
No comments yet — start the discussion below.