Mariusz Kulma · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22912827
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
A composite self-assessment score may predict correctness through task structure rather than discrimination of a model's own errors. We audit verbalized confidence minus 0.15 times the number of agents selected in a synthetic disagreement-identification task. In an initial free-cardinality dataset, score AUROCs were 0.992 and 0.994 for two local language models. Confidence was 1.0 in every analysed record and yielded AUROC 0.500. Replacing selected identities while preserving their count produced mean AUROCs of 0.977 and 0.995. A subsequent study used a different corpus with the required count supplied in the prompt. The primary analysis comprised five paraphrased-content cells, with 30 packets per cell shared by both models. Gemma reported confidence 1.0 in all 150 valid responses, including 16 errors. Phi reported 1.0 in 141 of 146 valid responses; 55 valid responses were incorrect. Within-cell AUROCs were 0.500 and 0.522; Phi's packet-bootstrap 95% interval was [0.481, 0.565]. Pooled score AUROCs nevertheless reached 0.626 and 0.675, compared with 0.626 and 0.660 for selection count alone. A computational audit reproduced the primary inference using the same data and random draws, and recovered inserted perfectly separating and constant confidence patterns. The studies differ in several design features and do not estimate a causal effect of fixing cardinality. They illustrate the need to distinguish composite-score discrimination, task competence and within-stratum error monitoring. Conclusions concern the observed prompts, configurations and synthetic task. Preprint and reproducibility artifact. AI-assisted preparation and critique are disclosed in the paper; no independent human peer review is claimed.
No comments yet — start the discussion below.