Ahmed Radwan, Athanasios V. Vasilakos, Shaina Raza · Preprints.org 2026 · 2026
DOI: 10.20944/preprints202609.2140.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
LLM-based agents increasingly combine model capabilities with memory, tools, and execution control to perform tasks across diverse environments. As research pursues more general intelligence, recent results highlight the importance of the agent harness: changing the execution system around a fixed model can alter performance and even reverse model rankings. Yet inconsistent reporting of harness configurations, resource budgets, and scoring procedures makes benchmark results difficult to interpret and compare. This survey reviews the development of agent harnesses and introduces a four-factor evaluation framework comprising the model, harness, environment, and evaluator, alongside the protocol under which they are tested. We map representative systems and benchmarks to this framework, examining evidence on task performance, reliability, safety, and human oversight. We distinguish what existing comparisons establish from effects that remain entangled, and provide factor-level reporting guidance to support consistent and reproducible evaluation. Finally, we identify research directions in adaptive harnesses, resource-aware comparisons, and the generalization of harness effects across models and environments.
No comments yet — start the discussion below.