When Scenes Shift: Transition-Referenced Adaptation for Vision-Language-Action Models
Abstract
Vision-language-action (VLA) policies can degrade substantially under scene shifts even when task-relevant state evolution retains a similar structure. In few-shot adaptation, target demonstrations directly supervise *what to do* in the new scene, but provide limited guidance on how shifted observations should reconnect to the policy's existing control representations. Motivated by these observations, we introduce **ShiftVLA**, a transition-referenced framework that treats the relation between internal control and task-relevant state evolution as a reference for few-shot scene adaptation. ***(i)*** To construct this reference, a transition learner captures demonstrated state progression through intermediate-state prediction and aligns transition relations with action-chunk similarity. ***(ii)*** A lightweight reader maps depth-aggregated Action-Expert control states into the transition space, while joint grounding shapes the underlying control representations with both transition and native action supervision. ***(iii)*** During target adaptation, source-calibrated control readouts drive depth-wise refinement, while gradient projection suppresses action-update components that conflict with the established correspondence. Under low-shot adaptation, **ShiftVLA** improves average success over the strongest baseline by 1.8 and 3.5 percentage points on LIBERO-Plus and RoboTwin 2.0, respectively, and improves real-robot task success by 7.0 percentage points under controlled scene shifts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.