Manabu Higashida · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22668213
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
University teaching and development lack a shared vocabulary for deciding which generative-AI model fits which use. This paper proposes user-grounded fidelity probes (here, probe means a diagnostic task whose answer the user already holds, not linear probing of internal representations): a diagnostic method that measures a model's capability on tasks for which the user already holds the ground truth, and reports it on graded scales rather than as a pass/fail verdict or a rank on a general benchmark. Responses to provocative claims are read qualitatively; the generation of running code is scored by execution and decomposed into two scores, reached depth (an ordinal level) and fidelity (decomposed in what follows into correctness, specification compliance, and process fidelity). In single trials, reached depth on nested-composition tasks appears to follow capability tiers; when the same tasks are re-measured with n = 5 repetitions in a controlled comparison that fixes provider and generation (the Claude three-tier set), the tier differences disappear, showing that single trials had rendered sampling variance as steps. What remains is not a staircase but a gradient of reach rate: only gpt-5 passes higher-order nesting robustly (5/5); the others sit in a high-variance region. Reached depth and reach rate are distinct evaluation facets (we do not claim statistical independence, and §6 notes that much of the reach-rate axis is a product of the bounded scale). We show, with an example from course development, how binary metrics render intermediate levels as discontinuous "emergence", and how the threshold decision sits on the user's side — a two-layer separation. The method builds on a companion paper and extends to exercise courses in general. English archival version (v1.0, 2026-09-08) of the Japanese manuscript posted on Jxiv (doi:10.51094/jxiv.5707). The data package (zip) is deposited with this record.
No comments yet — start the discussion below.