Seungho Lee · Journal of High School Science 2026 · 2026
DOI: 10.64336/001c.170179
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Hallucinations—fluent outputs containing incorrect or unsupported factual claims—remain an important obstacle to reliable use of large language models (LLMs). This study evaluates the scope and limits of HalluDetector, a reference-based detector that identifies contradiction-like evidence using lexical, numeric, unavailability, atomic-mismatch, and natural-language-inference rules. The evaluation comprises 650 prompt-reference entries, 7,800 archived first-line outputs, and a 1,850-row human audit spanning a legacy short-answer benchmark, a non-atomic pilot, and a balanced non-atomic add-on. Human agreement in the balanced audit was 96.8% for multiclass labels (Cohen’s κ = 0.866) and 100% for binary contradiction labels (κ = 1.000). The principal result is a substantial benchmark-transfer failure: detector precision fell from 84.6% on the legacy short-answer audit to 8.1% in the balanced non-atomic setting, where only 24 of 298 detector-positive outputs were judged to contain explicit contradictions. Weighted false-positive rates were highest for citation-grounded answering and nuanced uncertainty (18.6% each), compared with 0.9%–6.0% for the other evaluated categories. Error analysis showed that the detector frequently responded to false or unsupported propositions that were mentioned without endorsement, rejected, qualified, or discussed under source and evidence limitations. Thus, the principal failure was not simply recognition of proposition content but distinguishing that content from its assertion and epistemic status. Category-specific recall comparisons remain imprecise because contradiction-positive labels were sparse. These results show that strong contradiction-detection performance on compact reference-answer benchmarks can substantially overstate reliability in non-atomic language settings. HalluDetector is therefore best interpreted as a recall-oriented screen for explicit reference disagreement rather than a general detector of hallucination or semantic inadequacy.
No comments yet — start the discussion below.