WARD: Watching Actions via Latent-Space Reachability for Runtime Backdoor Defense in VLA Models
Abstract
Vision-language-action (VLA) models map visual observations and language instructions to robot actions, enabling a single policy to perform diverse manipulation tasks. However, backdoors implanted during training can preserve normal task performance while redirecting behavior when a specific trigger appears at deployment. In physical environments, such behavior can damage objects or equipment before task failure becomes apparent, and stopping afterward cannot undo the damage. Preventing these outcomes requires assessing actions before execution. This is challenging because a physically admissible action may still serve an unauthorized task, while failure detection alone does not establish whether a candidate action should be executed. We propose WARD, an external supervisor that combines predicted action effects with instruction-conditioned authorization to assess physical and task-consistency risk before actuation. WARD allows the candidate action or stops execution without modifying the deployed policy or accessing its internal representations. We evaluate WARD on OpenVLA-7B across four LIBERO suites and three backdoor attack families, with component ablations and physical-robot trials. WARD reduces suite-averaged progress-aware trajectory-matching attack success rates by 81.8–94.0 percentage points relative to no defense and achieves the highest aggregate defense utility among evaluated baselines, accounting for both attack suppression and clean-task preservation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.