Anuja Khatri · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23121190
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
At least six research programmes argue that language models alone cannot do reliable scientific reasoning. Each proposes something different: external verifiers, causal structure, neurosymbolic hybrids, world models, program synthesis, or entirely new architectures. Most of this work is theoretical or benchmark level. We report what happened when we surveyed all six and built a production system for cross domain scientific claim comparison on top of three of them. The system uses language models for extraction only. All judgment is deterministic code: typed subgraph isomorphism, a graded confidence ladder, and a versioned instrument capability map. A prior embedding based version produced 6,341 false positives and was killed. The rebuilt version, blind tested against the COGITATE adversarial collaboration, recovered the expert findings and surfaced additional ones. We document, for each programme, what we took, what we changed, what we left alone, and why. The theory belongs to the researchers we cite. The deployment and evaluation belong to us.
No comments yet — start the discussion below.