Elias Schlie · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23108425
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Accurately measuring fuzzy behavioral concepts like deception in large language models (LLMs) is notoriously difficult and many different methods have been proposed to do so. One strategy judges a model's response to a single user message, containing its role, motivation, context, and a question. This narrated framing differs from the interactive framing of deployment, where the model's role is established in a system prompt and the model directly addresses a counterparty. We ask how much of a model's measured deception rate is carried by this difference, decomposing the narrated-to-interactive transformation into two operations that can be varied independently: rewording the scenario from a single narrated block into a separate description and direct in-scenario user question, and splitting the delivery from a single user message into a system message plus a user message. Applying those operations to the narrated benchmark DeceptionBench (Huang et al., 2025), we evaluate their effect on deception rates across eight target models drawn from four vendor lineages. Both operations increase deception in seven of eight models, with the median rewording effect more than doubling the odds of deception and the median split effect more than tripling them. The operations compose roughly additively without amplifying one another. The consistency of these shifts across the four vendor lineages suggests that the framing a benchmark uses carries a meaningful share of the deception rate it reports, and that scores from differently-framed benchmarks cannot be read as comparable measures of the same underlying construct.
No comments yet — start the discussion below.