Cristian Ruvalcaba, Agentic AI Research Team · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.21247053
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
v0.3 (2026-10-03). Correction. An internal re-analysis of the released per-generation records (github.com/saluca-labs/cct) on 10 August 2026 found that the poisoned-reference results reported in v0.1 and v0.2 do not measure what the paper says they measure. The full re-analysis behind this correction, with the failure chain, a four-tier sensitivity analysis and the recomputation scripts, is published separately as A Substring Grader Over Truncated Chain-of-Thought Inverts a Pre-Registered Finding: A Null Result and a Post-Mortem (10.5281/zenodo.23125189). Publishing this correction took until now because the lab is small and it was queued behind other work. 20 of the 24 poisoned-reference generations and 14 of the 24 mindset-plus-poison generations reached the 1,200-token generation limit before producing a final answer. For those, the scorer read the last 120 characters of the reasoning trace and counted an “echo” whenever the planted wrong answer appeared there as a substring, including traces where the model was explicitly rejecting it. We therefore withdraw the claims that a poisoned reference is echoed as the final answer 46% of the time, that the mindset cuts this to 33%, that three poisoned trials flip back to the correct answer, and that the mindset demonstrably blunts the poisoned reference. The 4% echo reported for the correct-reference condition is the same substring artifact (the planted answer “1” matches the correct answer “12”) and is also withdrawn. What the records do support: a confident wrong reference made the model fail to finish within the limit in 20 of 24 runs, against 1 to 3 of 24 for the other conditions, and the mindset reduced that to 14 of 24. The accuracy results are unchanged (base 79%, nudge 79%, mindset 92%, correct reference 96%), but the mindset gain is three tasks of 24 and is not statistically significant at this sample size, so we no longer call the probe “powered”. This version also discloses that 6 of the 24 tasks were retrieval misses that served no mindset passages, fixes two internal inconsistencies (“19 reported here” and “+10”), narrows “structurally cannot commit” to the training-time failure, describes TKHR as patent pending, and fixes the author line. The paper is otherwise unchanged apart from punctuation. Cite the concept DOI, 10.5281/zenodo.21247053, which always resolves to the latest version. A position/response paper to the 2026 line of work showing that on-policy self-distillation (OPSD) degrades, rather than improves, the reasoning of small ‘thinking’ models by suppressing the deliberation tokens (‘Wait’, ‘Let’, ‘Maybe’) that carry multi-step search. We argue that for a large, under-named class of needs, temporary, task-scoped functional tuning, the right response is not to repair weight-level distillation but to avoid it. Curated Context Tuning (CCT). We describe serving a small model reasoning-shaped, human-readable, hash-chained curated context (‘mindsets’) on demand over the Model Context Protocol (MCP), leaving the base reasoning prior untouched. CCT does not commit the training-time failure the literature identifies: it never overwrites the prior, its corpora are δ_IT-shaped (method and heuristics) rather than δ_ref-shaped (answer keys), and its supervision is readable rather than silent. Empirical probe (n=144). An inference-time probe (24 reasoning-trap tasks × 6 context conditions, qwen3:32b, one sample per cell). A correct reference served as context does not reproduce the training-time collapse. A poisoned reference most often drove the model into reasoning that did not finish within the 1,200-token generation limit (20 of 24 runs, against 1 to 3 of 24 in the other conditions), so whether it would have adopted the planted answer is not measured. A δ_IT mindset lifted accuracy from 79% to 92% (three tasks of 24) at deliberation parity with base, while a generic ‘think-carefully’ nudge left it at 79%; the correct answer key reached 96% with deliberation below base. The direction matches the δ_IT and δ_ref distinction, but at this sample size none of the differences is statistically significant. CCT is positioned as complementary to (not a competitor of) weight-level methods: a curated mindset is precisely the clean, question-conditioned reasoning target such methods work to reconstruct, and it is auditable and attributable in a way weight-baked knowledge is not. This record is a preprint (v0.3); the probe harness, per-generation records, and transcripts are released for inspection.
No comments yet — start the discussion below.