
Gonzalo Emir Durante · Qeios 2026 · 2026
DOI: 10.32388/t420iv.4
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Hallucination-detection systems are usually judged by a single headline metric, computed under a benchmark design that determines what that metric can show. This paper’s contribution is not a detection algorithm — structural coherence scoring is established — but a methodology for measuring exactly what a benchmark proves and what it does not, applied to the SAS Multimetric Tribunal: a seven-module, GPU-free structural classifier combining lexical, entity-mutation, curvature, and logical-inversion signals via a cascading multiplicative penalty (\(\kappa_{D} = 0.56\), \(\kappa_{R} = 0.15\)). On a stratified 1,800-pair benchmark (HaluEval-Dialogue, HaluEval-QA, TruthfulQA), the tribunal reaches F1 = 98.99% with zero false positives under an A\(\rightarrow\)A identical-text sanity check — matched or exceeded by trivial baselines, including exact string matching (F1 = 100%). We report this explicitly: a benchmark that cannot fail on identical text cannot establish precision. Under a real negative control (R3: 127 valid paraphrase pairs), every method tested fails — the tribunal flags 126/127 as hallucinations (specificity 0.8%), matched by simple lexical baselines. No lexical-overlap method in our comparison distinguishes a valid paraphrase from a hallucination. We also report, as direct ablations rather than argument, that multiplicative combination outperforms weighted averaging of the same seven modules by 87.97 F1 points, and that SourceTargetGuard’s measured contribution is 3 of 1,800 true positives (0.17% of recall). The tribunal’s contribution is not a solved precision problem: it is an auditable, per-module architecture on which a semantic-equivalence guard can be added and measured — something a single-metric baseline offers no place to attach. This revision also reports a forensic audit distinguishing an earlier version’s undocumented-but-reproducible metrics from a stale sub-count, corrected in response to independent peer review.
No comments yet — start the discussion below.