Karel Hrubec · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23054200
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The Unmeasured Branch is a philosophical position paper about a second, often neglected role of AI evaluation. Evaluation does not only measure present systems. Once benchmark scores, preference signals, safety tests, cost targets, and other criteria influence training, model selection, scaling, and deployment, they also become developmental selection pressures shaping which properties future AI systems preserve. The paper argues that genuine progress in measured capabilities does not guarantee monotonic progress in poorly measured epistemic capacities. A successor model may become more accurate, faster, cheaper, more agentic, or more effective on benchmarked tasks while regressing in properties such as preservation of the original question, distinction between new evidence and new inference, recognition of underdetermination, calibrated claim ceilings, selective corrigibility, and the ability to stop when further reasoning adds no new evidential constraint. Rather than proposing another universal score, the paper argues for a multidimensional epistemic profile and predecessor–successor audits designed to detect regressions hidden by aggregate performance measures. The argument draws on contemporary work on benchmark validity, Goodhart effects, hallucination incentives, sycophancy, epistemic agency, and a longitudinal human–AI research programme. That programme is used as motivating case material, not as evidence that newer model generations are generally worse. The central developmental warning is simple: what AI developers choose to count as improvement can become part of the causal process determining what future AI becomes.
No comments yet — start the discussion below.