acceptodds
Under review as a conference paper at ICLR 2027

Action-Conditioned Predictive Consistency for Adversarial Detection in Vision-Language-Action Models

Abstract

Vision-Language-Action (VLA) models now let robots read a camera image and decide how to act. However, this also makes them vulnerable to visual perturbations. A carefully crafted change to a single frame can push the robot into executing a wrong action. On a real robot, this becomes a physical safety hazard, not merely a mislabeled image. Such frames must be detected before the policy acts on them. Existing defenses retrain or modify the policy itself, which is costly and alters a model that is already trusted. The closest learned monitor targets task failure, not adversarial attacks. What is missing is a runtime detector that runs alongside the policy and flags each attacked frame without modifying it. Our approach exploits the fact that control unfolds over time. At each step, the previous frame and the action the robot just took are already fixed before the adversary can perturb the current frame. The adversary can change the present, but it cannot affect this already committed past. We therefore predict the current frame's representation from the past alone, which gives a reference the perturbation cannot influence. Any attack strong enough to change the action must move the observed features away from this reference, and that discrepancy is our detection signal. We introduce PAC-CTCN, a per-frame detector that needs no change to the policy and comes with a provable detection guarantee. We prove that clean and adversarial frames separate once an attack displaces the features by more than twice the model's own prediction error. Both sides of this condition are measured directly rather than assumed. The guarantee is gray-box, and we state plainly that it is not robustness against an adversary who also knows the detector. We evaluate PAC-CTCN on three architecturally distinct VLAs across regimes of increasing difficulty. The detection condition holds with a clear margin throughout, and the detector outperforms both the closest learned monitor and standard detection baselines. It is a lightweight module, so it can drive a stop controller that turns detection into containment. Attacked runs then end in a safe halt at a small cost to clean-task success. Our source code will be made public after acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.