Iliya Angelov Radulov · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23143960
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Trigger-based attacks on robot-learning policies implicitly assume that the target model can perceive and use the channel in which the trigger is encoded. We examine this assumption through an attempted reproduction of a published text-triggered backdoor attack on a low-cost SO-101 robot using an ACT policy. The reproduction initially appeared to show a trigger- dependent failure, but the effect disappeared under a more controlled physical setup. Source- code inspection and a same-prompt control established that standard ACT has no language- conditioning pathway and that the apparent effect was instead explained by object-placement variability and contradictory action labels. We therefore formulate trigger transfer as a three- level question: whether the relevant pathway exists architecturally, whether it remains function- ally active after fine-tuning, and whether the training recipe preserves the required conditioning behavior. We then target the input modality ACT actually uses by training a visual back- door with freshly recorded episodes containing a distinctive background pattern. The visual trigger produces a consistent freeze-pose response on its native background while remaining distinguishable from behavior on a never-trained-on control background. The results show that shared hardware and tooling do not establish attack transferability: the target architecture, fine-tuning state, and training recipe must all be verified before a trigger-based claim can be interpreted causally.
No comments yet — start the discussion below.