SOMA SEMA: Action Space-Bound Attack that Awakens upon Benign Finetuning
Abstract
A vision-language-action (VLA) model is almost never deployed as released: before it can act, a practitioner finetunes it on its own robot, a step assumed safe because the finetuning data is the practitioner's own. That assumption fails. In SEMA, a provider distributes weights that come within a few points of a clean model's published task success wherever we benchmarked them and pass the screening we run, yet produce no unsafe behavior until the victim's own benign finetuning brings one out, with no poisoned example and no inference-time trigger. The same mechanism removes a hierarchical policy's safety guardrail and plants backdoors gated by a task, an object or a phrase. Whereas its text-domain predecessor carries a trigger that generalizes across the finetuning data, ours carries an action whose motor realization is bound to the action representation: it awakens on a different vendor's arm that shares the 7-dimensional single-arm space, and stays dormant both on a 14-dimensional bimanual robot and on the implant robot itself once the same actions are written in joint velocities or with permuted components. The semantic half of the guardrail payload is not so bound: a bystander's planner stops refusing though its arm never acts. Across four VLAs and two checkpoint types the attack reaches 74.0% to 86.5% of rollouts on the discrete-token models but 0.0% to 34.5% on the flow-matching ones, while the finetuned model stays within 8.5 points of a cleanly finetuned one on the victim's tasks. A short finetuning followed by screening then detects it on 58.0% of rollouts against 14.0% for a clean checkpoint. Severity throughout is a detector firing on the executed trajectory rather than a measured contact force.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.