acceptodds
Under review as a conference paper at ICLR 2027

When Images Read the Instruction: Bidirectional Attention in Vision-Language-Action Models

Abstract

Embodied reasoning and manipulation require rich interaction between visual, linguistic, and proprioceptive modalities, yet the dominant VLA backbone is a causal decoder that encodes the image before the instruction or robot state enters the context. Bidirectional attention is a natural remedy, as all tokens can freely attend to each other. Popular VLAs already adopt bidirectional observation encoding, but its effect has not been tested in a controlled comparison. In this work, we systematically study how bidirectional attention shapes representations for embodied tasks. We build matched paradigms from one causal Qwen3-VL-2B that vary in their embodied pretraining and attention mask: causal, flipped to bidirectional at policy time, or bidirectional and learned during pretraining. We show that only a learned attention mask improves manipulation over causal pretraining on the same data when using a final layer acton head. Activation patching, closed-loop instruction transplants, attention-path knockouts, and object-level attention measurements trace this difference. The attention mask decides where the instruction is represented: every bidirectional mask moves the instruction into the image tokens, and the policy reads its task from there. At what stage the mask is learned decides whether the instruction's attention stays on the task objects. A learned mask preserves the object selectivity that embodied pretraining creates, while a flipped mask does not. Our results suggest that bidirectional attention in VLA backbones should be aligned during pretraining rather than added at policy time.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.