Saluca Agentic AI Research Team · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22981455
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This version corrects citation errors found by an automated check and confirmed by hand. The force-aware re-sampling method FIRST was cited under the identifier 2606.12402, which is DIRECT, a paper on test-time compute routing; FIRST is arXiv:2606.12406, and the earlier statement that DIRECT's primary contribution is force-aware re-sampling has been removed (its compute-routing citations were correct and remain). DynaFLIP was cited under the identifier 2605.30864, which belongs to an unrelated cognitive science study; DynaFLIP is arXiv:2605.30350. This second error was found during the hand check rather than flagged automatically. The claims themselves are unchanged. This version has not had a full claim-by-claim audit. Vision-Language-Action (VLA) models have emerged as a dominant architecture for robot manipulation, yet their generalization failures are poorly understood at the mechanistic level. This paper synthesizes five specific findings from recent cs.RO preprints into a candidate structural reading of why VLA policies fail to generalize and what design interventions target those failure modes. The thesis, stated as a heuristic reading rather than a formal derivation, is: VLA policy generalization is jointly constrained by the quality of the perceptual substrate fed into the policy, the temporal granularity at which world and action predictions are coupled, the distribution of training data relative to contact-critical moments, the granularity of language supervision, and the architecture-specific signatures of motor-command failure. Each constraint operates at a distinct stage of the policy pipeline, perception, temporal planning, data curation, language conditioning, and runtime monitoring, and the five together suggest that improving any single stage while ignoring the others yields diminishing returns. Corpus sources span cs.RO primary-category preprints covering dynamics-aware visual pre-training arXiv:2605.30350, asynchronous world-action temporal decoupling arXiv:2606.09811, noise-dependent suboptimal data usage arXiv:2606.12365, fine-grained language supervision arXiv:2605.27284, and architecture-matched action monitoring arXiv:2605.28726. Supporting evidence on force-aware data re-sampling arXiv:2606.12406, test-time compute routing arXiv:2606.12402, speed-controllable execution arXiv:2606.06491, and attention-guided safety filtering arXiv:2606.09749 is incorporated as corroborating or weakly-connected addenda where noted. The primary falsification path is a controlled ablation on a shared benchmark (e.g., RoboTwin or LIBERO) that independently removes each of the five interventions while holding the other four fixed, measuring the marginal generalization drop per intervention; if the drops are not additive, the independence assumption underlying this heuristic reading is falsified. Authorship: Saluca Agentic AI Research Team (Saluca LLC). AI-drafted synthesis from an arXiv preprint corpus, originally drafted 2026-06-16, produced under the direction of Cristian Ruvalcaba, the accountable human author. Not peer-reviewed. AI disclosure. This work was produced with an agentic AI research apparatus operated by Saluca Labs. The apparatus drafted, searched and analysed under direction. Cristian Ruvalcaba is the human author and is accountable for the content. No AI system is listed as an author or contributor, because authorship entails accountability that a model cannot hold; this disclosure is the credit, and it is deliberately the whole of it.
No comments yet — start the discussion below.