Pranay M. Mahendrakar · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23168093
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0. Two bodies of work describe the same deployed language models in opposite terms. One reports that models behind a fixed product name change over time and calls for continuous monitoring. The other reports that a single model, queried twice under settings meant to be deterministic, returns different answers, and that benchmark scores move by several points under seeds, hardware, batch size and prompt format alone. Read together, they raise a methodological question that neither answers on its own: when is an observed difference between two dates evidence that the provider changed something, rather than a draw from the noise of the measurement? This paper argues that the common framing of that question - drift claims ignore the noise floor - is too strong in one direction and too weak in another. It is too strong because the most careful longitudinal studies did estimate a floor: the original ChatGPT drift study measured same-version disagreement on one task, a daily study of an unpinned endpoint ran repeated queries and z-tests, and a clinical study separated same-day from across-day variation. It is too weak because a noise floor is not a property of the model that can simply be measured; it is a declaration of what counts as "the same", and published monitors declare three different things. Tests that treat any change in the output distribution as change detect hardware swaps and serving-stack updates readily; tests of task capability need far larger samples and are rarely run. The confirmed changes located are mostly announced re-pointings of a name or provider infrastructure events. A capability change under a pinned identifier, with no announced event and a calibrated floor, was not established by any study found. The paper separates the three claims, collects the published magnitudes of change and of noise in one table and one figure, states the evidence standard as an algorithm, and lists the studies that would settle the rest.
No comments yet — start the discussion below.