Danylo Yefimov · Věda a perspektivy 2026 · 2026
DOI: 10.52058/2695-1592-2026-9(64)-539-549
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The article is devoted to the systematization of approaches to the evaluation of LLM-based autonomous agents.It is shown that the transition from single-turn language models to autonomous agents fundamentally complicates the problem of quality assessment, since a correct final result is no longer sufficient: the execution trajectory itselfthe choice of tools, the validity of their arguments, the ordering of actions, the absence of redundant steps, and the ability to recover from failuresdetermines whether the task was solved reliably.The paper describes the principal groups of behavioural and outcome-oriented metrics, including tool selection accuracy, argument correctness, action sequencing, redundancy rate, error recovery, as well as helpfulness, correctness, groundedness, efficiency, and safety.Particular attention is given to the stochastic nature of large language models, which makes the assessment of average performance insufficient and necessitates the analysis of reproducibility, stability, and statistical significance of the obtained results.The corresponding metrics are considered, including Pass@k and pass^k, the sample variance and standard deviation as measures of run-to-run variability, and confidence intervals together with paired significance tests for the comparison of models.The use of large language models as evaluators (LLM-as-a-Judge) is examined separately, along with chance-corrected agreement measures (Cohen's and Fleiss' kappa, Krippendorff's alpha) and the systematic biases inherent to such evaluatorsposition, verbosity, and self-preference biasas well as strategies for their mitigation.The article further analyses public benchmarks for the main classes of agentsconversational, coding, research, and computer-use systemsand substantiates the limited reliability of benchmark results owing to contamination, saturation, and reward Věda a perspektivy № 9( 64) 2026
No comments yet — start the discussion below.