igauna Errapel Iñaki Gauna Leon · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22981851
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
SummaryIn a single conversation, the model produced several fluent statements that were unsupported orinconsistent with what it had said earlier. The user caught each one; the model recognised eacherror immediately once it was pointed out. What was missing was a review step, not knowledge.From this the conversation produced three proposals:• A. A “Verify claims” control, independent of the effort level (Off / Auto / Always).Reasoning longer is not the same as checking claims.• B. Per-claim status labels (verified, corrected, not verified, not verifiable), so users can seewhere to apply their own scrutiny.• C. An expert evaluation channel where domain experts assess unaltered model outputblind, and their detections are compared with an automatic evaluator to expose the model’sshared blind spots.The common thread: today the burden of detecting confabulation falls on the user, who usuallycannot tell a confabulated answer from a correct one because both arrive with the same speedand fluency. These proposals aim to make failures visible and predictable, not to promise thatthey disappear
No comments yet — start the discussion below.