acceptodds
Under review as a conference paper at ICLR 2027

LOOK: Latent Knowledge Is Overlooked, Not Overwritten, in Vision-Language-Action Models

Abstract

Failures of vision-language-action (VLA) models on concepts unseen in robot demonstrations are often attributed to semantic damage during action training.We find that target knowledge can be weakened yet remain useful for control: the surviving signal can be overlooked by the action pathway. We study this possibility with CrossBind, which measures target identification and closed-loop behavior separately. Across three action-pretrained backbones, an unfitted likelihood readout identifies placement targets on 0.56 to 0.69 of episodes against a 1/3 chance level. Yet same-checkpoint pairing in two task-adapted policies reveals only weak associations between correct identification and correct placement. This discrepancy motivates a semantics-to-action access bottleneck. RECALL tests this account by giving the retained signal an explicit path to action. With the backbone frozen, demonstration target indices and the original action loss train the action expert and interface; the inherited readout selects targets at inference. At fixed policy weights, replacing a location-only condition with readout-selected targets raises success from 0.215 to 0.438 on π0.5, from 0.056 to 0.722 on InstructVLA,and from 0.049 to 0.514 on X-VLA. These results show why preserving semantics must be accompanied by attention to its use: retained target knowledge can support behavior when given an effective route to action

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.