acceptodds
Under review as a conference paper at ICLR 2027

Correctable Means Corruptible: The Reasoning Chain of Vision-Language-Action Policies Under Adaptive Attack

Abstract

Vision-language-action policies map camera images and natural-language instructions to a robot's motor actions. Some of these policies reason before they act, writing a reasoning chain in text and decoding their actions conditioned on it. That chain is offered as an interface that a person can read to understand the robot's behavior and rewrite when the behavior is wrong. An attacker who reaches the chain can rewrite it too. We therefore ask whether the chain can be exposed as an interface and still be secured by monitoring it. We answer with two measurements on DeepThinkVLA and CoTinyVLA, two such policies that write their chains at different times, in four simulated task suites of the LIBERO benchmark. In the first measurement, a chain corrupted by renaming its objects lowers task success on every suite of both policies, even with an intact instruction. Conversely, injecting the clean chain with a corrupted instruction recovers part of what that instruction cost. On both policies, the chain that corrects the robot can also corrupt it. In the second measurement, two detectors that read only the chain catch the corrupted chains when the attacker does not know them. Against an adaptive attacker that knows what they check, neither detector scores above chance on either policy. In every case the attack keeps the corruption's harm to task success. Yet on every feature the detectors check, the attacked chains look at least as clean as the clean chains. We call this outcome disguised corruption. We argue that detection over features read from the reasoning chain fails as a class whenever an attacker can achieve disguised corruption against the features it reads. This argument instead points to a detector that reads a channel the attacker does not write. We therefore pre-register an action-only task-consistency detector with an attacker built against it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.