David Howard · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23091239
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Description This deposit comprises a two-paper case study and supporting evidentiary appendix documenting a previously unnamed AI reasoning failure mode: framework-selection failure. The investigation began when queries about P1 and P2 stars in globular clusters — using Arabic numerals consistent with the intra-cluster classification — were posed to ChatGPT. The model immediately substituted Roman numerals, responding as though the galactic-scale Population I/II framework had been invoked. No ambiguity was flagged. The subsequent reasoning was internally coherent and factually accurate — within the wrong framework. Paper 1 (Narrative) presents the case study as an accessible essay. It introduces a four-stage reasoning model in which framework selection is Stage 1, prior to fact retrieval, reasoning and conclusion. It argues that failures at Stage 1 produce coherent but contextually incorrect outputs that do not announce themselves as errors — in both humans and AI systems. It distinguishes framework-selection failure from hallucination and cognitive bias, and draws implications for organizational governance of LLM use with specialized or proprietary terminology. Paper 2 (Method) presents the same investigation as a structured scientific report, following the standard observe–question–hypothesis–experiment–results–analysis–conclusion format. It documents the exact queries used, the model's responses verbatim, and the conditions under which the failure was detected and corrected. Appendix A reproduces the original ChatGPT session screenshots (Screenshots 1–3, June 10, 2026) and the warning sidebar from Jan Hattenbach's "Ancient Star Polluters," Sky & Telescope, July 2026 (Screenshot 4), which independently confirms that the framework ambiguity is recognized in the professional literature. The production of this work is itself a demonstration of its argument: ChatGPT served as test subject and independent critic, Google Search as factual verifier and Claude (Anthropic) as synthesizer and editor — all directed by human judgment throughout.
No comments yet — start the discussion below.