Artem Orlovskyi, Gennadiy Kyselov · Інфокомунікаційні та комп’ютерні технології 2026 · 2026
DOI: 10.36994/2788-5518-2026-01-11-28
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The rapid evolution and adoption of large language models in domains with a high cost of error, including healthcare, law, finance, and education, where incorrect or factually inaccurate responses may lead to erroneous decisions, necessitate the study of the theoretical and practical aspects of the interpretation, calibration, evaluation, and verification of large language model outputs.This article systematizes contemporary approaches to the interpretation and verification of large language model decisions and identifies the roles of trust metrics in constructing protocols for assessing their reliability. The study analyzes methods of mechanistic interpretability, including sparse autoencoders and computational circuit tracing, approaches to chain-of-thought analysis, calibration metrics such as Expected Calibration Error (ECE) and Brier score, semantic entropy, verbalized confidence, and conformal prediction. Particular attention is devoted to the standardized evaluation frameworks HELM and LM Evaluation Harness, the LLM-as-a-judge paradigm, as well as factuality assessment methods including FActScore, SAFE, and SelfCheckGPT. Based on the conducted analysis, a five-level taxonomy of large language model interpretation methods is proposed, covering the range from the analysis of individual neuron activations to behavioral testing at the level of complete textual outputs. Trust metrics are grouped according to their level of practical maturity, while factuality verification approaches are compared in terms of verification granularity, dependence on external knowledge sources, and computational cost. It is demonstrated that proper probability calibration is a necessary prerequisite for the practical application of interpretation and verification results. Future research should focus on the development of comprehensive evaluation protocols covering the entire lifecycle of a large language model – from training to deployment in production systems.
No comments yet — start the discussion below.