Momen Ghazouani · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23124587
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
LinearBench is a framework for evaluating language models on multi-round build-and-refine tasks, in which an artifact is revised under feedback and judged against a hidden, weighted acceptance specification. A deterministic user simulator supplies feedback at four levels of informativeness, from generic self-review to machine-readable diagnostics, which separates a model's capacity for self-correction from its dependence on external guidance. The paper defines trajectory-level metrics, including stable turns-to-acceptance under right censoring, a regression and churn decomposition, a headroom-weighted feedback efficiency, a parameter-free persistence-based Convergence Score, and a Self-Sufficiency Ratio, and proves that summaries computed from the quality sequence alone cannot detect regressions. A synthetic study with parameterized model profiles illustrates cases in which conventional and proposed summaries rank the same profiles differently. This is a measurement proposal that reports no evaluation of existing language models; it specifies a pilot design, falsifiable hypotheses, and open problems, and invites collaboration on task construction, verification, and human validation of the simulator.
No comments yet — start the discussion below.