Reflective VLA: In-Context Action Consequences Make VLAs Generalize
Abstract
Most vision-language-action (VLA) models are reactive: they predict actions from the current instruction and observation. However, deployment-specific factors such as camera geometry, robot calibration, and actuation bias can be difficult to identify from a single observation, limiting the policy's ability to adapt its behavior based on observed action consequences. We propose Reflective VLA, which conditions each decision on observation–action–consequence triplets collected during the policy's own task execution. Each triplet records what the robot observed and executed, and how the scene changed afterward, providing action-aligned feedback for subsequent decisions without test-time parameter updates. Architecturally, Reflective VLA routes all observation modalities through the VLM under shared attention, allowing the action expert to attend directly to past triplets and the current observation. A block-causal mask enables parallel multi-frame training without leakage and supports KV-cached inference. Reflective VLA maintains strong performance on LIBERO and improves success on SimplerEnv-Bridge. Under deployment perturbations on LIBERO-Plus and LIBERO-Plus-Hard, it improves average success by 5.4 and 4.2 percentage points over a matched reactive baseline, respectively. Hard evaluates deployment configurations outside the task-specific training distribution within camera and calibration perturbation families represented during training. Controlled ablations support the joint use of historical actions and observed consequences beyond visual history alone, and real-robot studies provide complementary evidence under camera shifts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.