Nicholas Kasdaglis · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.20820156
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
A language model can be wrong and confident simultaneously, and the confident case matters for deployment: fluent, assertive, and invisible to signals designed to screen low-confidence output. We identify a structural reason for this blindness while showing that the model still retains internally accessible information about whether its answer is correct. We prove an input-invariance theorem for an idealized random-matrix transformer model: when input-dependent gating affects a sublinear number of dimensions and has bounded operator norm, the normalized per-layer transition-Jacobian participation ratio converges to an input-independent constant. Its endpoint-blindness corollary implies that unsupervised macroscopic geometric or spectral readers see the class-dependent part of this geometry diluted by model width below the measurement noise floor, leaving them operationally at chance, with the dilution deepening as models widen. On a nonlinear network, participation-ratio medians differ by 0.87% (AUROC 0.531). Correspondingly, semantic entropy detects genuine uncertainty (AUROC 0.774) but is near chance on adversarial confident error (0.476–0.563 across twelve models and four families). Crucially, blindness is not absence of information: our entity-grouped, out-of-fold commitment-band probe recovers a truth-related direction from frozen last-token residual states, detecting confident error at AUROC 0.731–0.849 across the powered model ladder and retaining AUROC 0.720 [0.674, 0.762] specifically on errors semantic entropy misses. We convert this direction into a Governor that identifies the error regime, routes a regime-matched correction, and refuses when residual risk remains high. Routing is necessary because the wrong regime's fix actively harms. Across five models and three families, routing beats the best single fix by +4.7 to +10.5 points, with every bootstrap interval excluding zero and McNemar p < 0.001 throughout; routing every item to its regime-matched corrector improves accuracy by +13.0 to +20.4 points over base. Intervention analysis reveals an asymmetric subspace where corruption survives orthogonalization while rescue collapses. Together, these results expose a separation between outward confidence, coarse internal geometry, and recoverable correctness information, showing that confident errors can be diagnosed and acted on rather than treated as terminal failures.
No comments yet — start the discussion below.