J. Marco · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23101791
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Large language models can now read and compare scientific papers at a scale that is inaccessible to individual researchers, but scientific corpora are normally used as if all statements carried comparable epistemic weight. We propose a simple alternative: let an LLM decompose the literature of a field into atomic claims, attach each claim to its supporting evidence and citation context, assign a local confidence calibrated against an explicitly defined outcome, and propagate confidence through a signed graph in which only credible evidence transmits support or contradiction and correlated sources are counted once. Local scores are assigned by tiered adjudication: two LLMs from different model families score each claim independently, a third model arbitrates when they disagree, and claims that remain uncertain, or whose errors would propagate furthest, are escalated to human referees. Paper-level confidence is then obtained only as an aggregate of claim-level scores. The resulting corpus can be used for retrieval, for confidence-weighted continued training, or for training conditioned on explicit epistemic labels, so that disputed claims are retained as disputed rather than deleted. The procedure can be iterated: a model reviews the literature, produces an evidence graph, is retrained from a fixed base using the annotated corpus, and then reviews the corpus again. We outline safeguards against circular self-reinforcement and distributional thinning, and propose a minimal experiment combining replication-outcome datasets with scientific claim-verification benchmarks. The central hypothesis is testable: recursive epistemic annotation should improve closed-book calibration on scientific claims compared with uniform training on the same literature.
No comments yet — start the discussion below.