Vladyslav Grybennikov · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23005704
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Classical software testing separates end-to-end, integration, and unit evidence, whereas prominent agent benchmarks aggregate task outcomes. For tool-using AI agents, an outcome alone cannot distinguish a correct state reached through an invalid path from a wrong state reached after plausible intermediate decisions. We implemented a trace-based evaluation harness that tests final environment state, execution-trace conformance, and frozen decision-component fixtures. Each task defines a goal state, an expected dependency graph, and prohibited patterns, while instrumentation records tool calls and mock-database snapshots. We evaluated a stateful telecom agent using deterministic defect specimens and live executions from Qwen and Claude Opus configurations. The specimens produced disagreement in both directions, showing that Levels 1 and 2 do not subsume one another, with Level 3 providing separate component evidence. Across 27 episodes per configuration, end-to-end pass1 was 0.444 for Qwen and 0.926 for Claude Opus, while fixed-rotation pass3 was 0.444 and 0.889. A cancellation trace exposed a policy-retrieval-to-action gap: the relevant policy and pending work appeared in the trace, but required cleanup operations were omitted while frozen component fixtures remained clean. Judge–oracle agreement was 0.444 for Qwen and 0.852 for Claude Opus. Layered testing therefore complements task success with evidence that separates outcome failure, path divergence, and isolated component error.
No comments yet — start the discussion below.