acceptodds
Under review as a conference paper at ICLR 2027

Reflective VLA: In-Context Action Consequences Make VLAs Generalize

Abstract

Most vision-language-action (VLA) models are reactive: they predict actions from the current instruction and observation. However, deployment-specific factors such as camera geometry, robot calibration, and actuation bias can be difficult to identify from a single observation, limiting the policy's ability to adapt its behavior based on observed action consequences. We propose Reflective VLA, which conditions each decision on observation–action–consequence triplets collected during the policy's own task execution. Each triplet records what the robot observed and executed, and how the scene changed afterward, providing action-aligned feedback for subsequent decisions without test-time parameter updates. Architecturally, Reflective VLA routes all observation modalities through the VLM under shared attention, allowing the action expert to attend directly to past triplets and the current observation. A block-causal mask enables parallel multi-frame training without leakage and supports KV-cached inference. Reflective VLA maintains strong performance on LIBERO and improves success on SimplerEnv-Bridge. Under deployment perturbations on LIBERO-Plus and LIBERO-Plus-Hard, it improves average success by 5.4 and 4.2 percentage points over a matched reactive baseline, respectively. Hard evaluates deployment configurations outside the task-specific training distribution within camera and calibration perturbation families represented during training. Controlled ablations support the joint use of historical actions and observed consequences beyond visual history alone, and real-robot studies provide complementary evidence under camera shifts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.