acceptodds
Under review as a conference paper at ICLR 2027

The Repair Must Reach the Backbone: A Controlled Comparison of Instruction-Following Fixes for Vision-Language-Action Policies

Abstract

A vision-language-action policy can clear every task in its training suite and still reach for the object its instruction did not name. Repairs exist, and each one arrives on its own base policy, its own demonstrations and its own benchmark, so the field cannot say which ingredient does the work. We hold the base, the demonstrations and the protocol fixed, and compare sixteen arms on one ruler: change one word of the instruction, hold the scene and the robot pose, and record which object the arm reaches, against a random-motion floor. The base reaches the named object in 36% of trials. A swap-differential probe shows that its backbone already picks out the newly named object, so the natural repair is to teach the action expert to read what the backbone computes. That repair does nothing. Training the action expert alone, at either rank, leaves the named-object rate where the base had it, and on the decorrelated demonstrations it also cuts success on the trained tasks by more than half. Given only what the robot observes, no repair that leaves the vision-language backbone frozen passes 53%. Every repair that passes it from the robot's own observations reaches the backbone, and the two that lead decorrelate the instruction from the scene layout: counterfactual relabelling, which edits the word, and low-rank adaptation on position-randomized demonstrations, which moves the object. The second reaches 97.7% across three seeds. Activation patching locates that access: in the pairs the base itself tells apart, swapping the keys and values of six consecutive backbone layers flips its choice 29 times in 32, no other band flips more than once, and the repair keeps that band and makes it more reliable. The recipe then carries off the checkpoint it was tuned on. A second base rises from the same 36% to 77%, a third-party counterfactual benchmark shows the repaired policy following the instruction where the base replays its demonstration, and objects whose name and mesh never occur in training are selected at twice the chance rate. We release the policies, the demonstrations and their generator, and the measurement code.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.