Anastasia Goudy · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22781578
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Background. Conversational AI is increasingly used to interpret evidence, troubleshoot problems, make recommendations, and evaluate competing claims. In sustained exchanges, an assistant may establish a substantive position that later encounters conflicting information from outside the current user-AI conversation. Existing work on trust, reliance, sycophancy, persuasion, and belief change does not directly measure this source-conflict event. Objective. This study developed and evaluated the naturalistic inter-rater reliability of the Source-Conflict Episode (SCE), defined by four required elements: a prior assistant commitment, an attributable outside challenge, material conflict, and subsequent assistant uptake. Methods. SCE measurement was developed through staged synthetic and naturalistic studies separated by prospective freezes. Integrity audits showed that early synthetic materials contained class cues, and a held-out WildChat study showed that near-perfect overall agreement could remain uninformative when positive cases were absent. Retrieval and construct coding were therefore separated. After SCE Codebook v2.2 FINAL and Scanner V1 were frozen, blind cross-corpus feasibility identified ShareChat as the first corpus to fill the prespecified candidate cap. A deterministic enriched reservoir of 290 authentic conversations supported sequential blinded coding by Claude Sonnet and OpenAI GPT-5.6 Sol (High reasoning). An administration-order mismatch was discovered from identifiers before outcomes were inspected. The reconciled 48-item checkpoint remained the primary analysis; outcome-blind completion of the already-administered union produced an expanded paired sample of 87 conversations. Results. In the primary checkpoint (N=48), raw agreement was .750 (item-bootstrap 95% CI .625 to .855), positive specific agreement (PSA) .600 (.357 to .800), negative specific agreement (NSA) .818 (.702 to .907), Gwet AC1 .562 (.309 to .778), and nominal Krippendorff alpha .424 (.132 to .689). The prespecified positive-information gate was met with 21 positive-union cases and 9 joint positives. In the outcome-blind expanded sample (N=87), raw agreement was .816, PSA .724, NSA .862, AC1 .669 (95% CI .509 to .813), and alpha .589 (95% CI .400 to .747), with 21 joint positives and 37 positive-union cases. Q1 discordance was asymmetric: 15 conversations were Sol-positive/Claude-negative and one showed the reverse pattern (exact two-sided McNemar p=.000519). Post hoc review located most discordance near the task-fidelity, attributable-source, and material-conflict boundaries. Conclusions. SCE v2.2 yielded a reproducible positive core in authentic conversations, with moderate primary reliability and somewhat higher agreement in the outcome-blind expanded sample. The study provides an initial naturalistic reliability basis for this event-level measure; human criterion validation and gate-level reliability remain the next measurement steps. Keywords: human-AI interaction; source-conflict episode; epistemic arbitration; inter-rater reliability; conversational AI; sycophancy; LLM-based coding; measurement development; AI safety
No comments yet — start the discussion below.