Shubham Jha · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22967544
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
In multi-step LLM reasoning, an early unsupported step can propagate to the final answer. We propose Project Sphere, a search procedure that combines hallucination signals at the level of individual reasoning steps. Sphere expands a reasoning graph breadth-first, scores every node by a weighted combination of model confidence, knowledge-grounded entailment, and cross-path convergence weighted for source diversity, and deepens promising paths best-first. When a node's risk exceeds a threshold, Sphere spawns a bounded sub-sphere that resolves the doubtful step before merging its result back. We formalize the risk scores, give the algorithm with cost and termination analysis, show why correlated samples cap the value of agreement, and specify an evaluation protocol with four benchmarks, seven baselines, and ablations.
No comments yet — start the discussion below.