Jie Meng, Jin Mao, Mingchang Ma, Gang Li · Information Processing & Management 2026 · 2026
DOI: 10.1016/j.ipm.2026.105154
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Large language models (LLMs) have become effective evaluators across a wide range of tasks, but in rigorous contexts such as educational measurement, a model that can produce plausible scores is not yet a reliable rater. A central limitation of current LLM evaluators is that their scoring logic remains implicit and is not consistently applied. We propose a rule-augmented evaluation framework that turns latent scoring logic into explicit, learnable scoring rules, decomposing reliable evaluation into two stages: rule discovery , which learns human-aligned scoring rules from labeled data through an LLM-assisted Monte Carlo Tree Search (MCTS), and rule execution , which applies these rules either through training-free prompting (Chain-of-Rule, CoR) or through reinforcement learning that trains rule-grounded reasoning into the evaluator (Rule-augmented Evaluator, RuAE). Across four evaluation tasks (essay scoring, scientific relevance ranking, review rating, and summarization) and both open- and closed-source backbones, rule augmentation yields the clearest gains on tasks requiring multi-aspect reasoning, including essay scoring and scientific relevance assessment (e.g., on essay scoring, RuAE raises quadratic weighted kappa from 0.286 to 0.379), whereas its gains on the lower-dimensional Amazon rating task are limited. Beyond predictive accuracy, we assess the evaluator as a rater through a psychometric diagnostic layer based on the Many-Facet Rasch and Partial Credit models. The diagnostics reveal task-dependent dimension-separation patterns; on SummEval, where a direct comparison is available, RuAE lowers the average inter-dimensional correlation from 0.886 under CoR to 0.800. They also show lower middle-score concentration under RuAE than CoR and characterize task-specific estimated information coverage across latent-quality regions. These results indicate that making scoring logic explicit and learnable moves LLM evaluators a step closer to inspectable rating systems.
No comments yet — start the discussion below.