When Images Read the Instruction: Bidirectional Attention in Vision-Language-Action Models
Abstract
Embodied reasoning and manipulation require rich interaction between visual, linguistic, and proprioceptive modalities, yet the dominant VLA backbone is a causal decoder that encodes the image before the instruction or robot state enters the context. Bidirectional attention is a natural remedy, as all tokens can freely attend to each other. Popular VLAs already adopt bidirectional observation encoding, but its effect has not been tested in a controlled comparison. In this work, we systematically study how bidirectional attention shapes representations for embodied tasks. We build matched paradigms from one causal Qwen3-VL-2B that vary in their embodied pretraining and attention mask: causal, flipped to bidirectional at policy time, or bidirectional and learned during pretraining. We show that only a learned attention mask improves manipulation over causal pretraining on the same data when using a final layer acton head. Activation patching, closed-loop instruction transplants, attention-path knockouts, and object-level attention measurements trace this difference. The attention mask decides where the instruction is represented: every bidirectional mask moves the instruction into the image tokens, and the policy reads its task from there. At what stage the mask is learned decides whether the instruction's attention stays on the task objects. A learned mask preserves the object selectivity that embodied pretraining creates, while a flipped mask does not. Our results suggest that bidirectional attention in VLA backbones should be aligned during pretraining rather than added at policy time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.