Dian Yu · Figshare 2026 · 2026
DOI: 10.6084/m9.figshare.34063563.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Stateful neural sequence models increasingly trade tokenwise recurrence for chunk-parallel execution. In Test-Time Training (TTT) layers, however, chunk boundaries determine fast-weight reference states, local rotary phases, and learned inner-learning-rate indices. Natural text supplies no such boundaries, creating a deployment robustness problem that aggregate accuracy can hide. We show that moving only the internal chunk-grid origin changes the predicted token on 16.15% and 15.63% of 384 held-out LAMBADA examples from frozen 350M and 760M TTT-MLP checkpoints. We propose Polyphase TTT, a training-free inference mechanism that treats chunk origin as a nuisance coordinate. It evaluates internally consistent placements of the same trained-width grid, retains the official placement as the prediction anchor, and admits only a bounded residual from the alternatives. A confidence cascade evaluates additional placements only for uncertain examples. Prediction-flip rates fall to 11.72% and 10.94%, removing 17 and 18 paired instability events while introducing none. Canonical accuracy is preserved at 350M and changes descriptively from 42.97% to 43.49% at 760M. On a disjoint 243-question SQuAD and HotpotQA transfer check, five prediction-instability events are removed and none introduced while canonical accuracy remains unchanged. The cascade averages 1.39–1.44 evaluations, incurs 1.51–1.65× reference latency, and adds less than 0.4% peak allocated memory. Polyphase TTT changes neither checkpoint weights nor fast-state updates, providing a practical route to more stable inference from released models without retraining.
No comments yet — start the discussion below.