Saleh Aldaajeh · OSF Preprints (OSF Preprints) 2026 · 2026
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
When several AI language models talk a question over and settle on an answer, it is tempting to treat their agreement as a sign that the answer is right. This project asks whether that trust is earned. On hard factual questions where the models' first answers disagree, the majorities they reach after discussion are wrong about three times in four. Majorities built from their independent first answers are wrong less than a quarter of the time. Discussion makes the models more accurate on average, yet it drains their agreement of its value as evidence. This repository holds the code, pre-registrations, run records and analysis scripts behind the paper. Author names are withheld during double-blind review.
No comments yet — start the discussion below.