Aradhana Rai, Vijay K. Madisetti · ACM AI Letters 2026 · 2026
DOI: 10.1145/3847307
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Multi-agent LLM systems are widely deployed under three implicit assumptions: that self-confidence calibration transfers to peer evaluation, that errors made by distinct instances are statistically independent, and that any sufficiently calibrated model is an equivalent verifier. We test all three across GPT-4o, Claude Sonnet 4.6, and Llama-3.3-70B. Self-confidence and peer-resistance dissociate: Claude shows the strongest self-confidence signal yet accepts every deliberately-wrong peer output, the opposite of what unidimensional calibration predicts. Four independent GPT-4o instances agree on the same wrong answer \(56\%\) of the time on obscure factual queries versus \(0.4\%\) predicted under independence—a \(140\times\) violation. A pre-registered cross-family replication ( \(N=200\) , three architecturally distinct families) yields \(55\%\) three-way agreement, statistically indistinguishable from the same-family baseline (one-sided binomial \(p=0.41\) ): family diversification does not reduce correlated hallucination. Verifier choice produces a 42-percentage-point spread in error propagation. We propose a three-phase framework treating self-confidence assessment, peer-resistance capability, and error correlation as orthogonal axes requiring distinct training signals, and identify query-inherent rarity, not shared training, as the dominant driver of correlated failure.
No comments yet — start the discussion below.