Thinking More, Moving Less: Stealthy Availability Attacks on Reasoning VLA Models
Abstract
Reasoning vision-language-action (VLA) models are increasingly studied for physical AI, yet the robustness of their computational efficiency during execution remains largely unexplored. We investigate a novel behavior-preserving availability attack, in which a static visual patch increases both the number of policy calls and the inference cost per call while preserving task-directed behavior. We introduce SPADE, a two-stage framework that separates the discovery of executable slow behavior from its realization through pixels. First, a learnable embedding suffix elicits smaller task-directed movements and longer reasoning from a frozen VLA, and closed-loop validation retains only behaviors that increase both cost factors while satisfying task, motion, and action-validity constraints. Second, a visual patch distills these behaviors by matching reasoning-exit decisions and actions, with supervision continually refreshed on the patched policy's own prefixes and visited states. At deployment, only the fixed patch remains, leaving the model weights, instructions, decoding procedure, and controller unchanged. Across three reasoning VLAs and two robotic manipulation benchmarks, SPADE substantially increases cumulative inference time while largely preserving task success and trajectory geometry, and achieves greater cost amplification than direct patch optimization under the same optimization budget. These results reveal an availability vulnerability that conventional task-success evaluation can overlook.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.