acceptodds
Under review as a conference paper at ICLR 2027

In-Context Structured Demonstration enables Generalizable Vision-Language-Action Models

Abstract

Vision-language-action (VLA) models often degrade sharply in out-of-distribution (OOD) scenarios. In-context imitation learning (ICIL) enables test-time adaptation by retrieving expert demonstrations and appending them as contextual prompts. Existing frameworks, however, underutilize these trajectories: by conditioning on raw observations and continuous actions alone, the policy observes what the expert did without the task structure that explains why. We propose StellaVLA, a framework that advances ICIL by distilling in-context structured rationales into VLA models. First, an offline pipeline extracts a hierarchical rationale from raw expert demonstrations without manual annotation, articulating both high-level semantic sub-goals and fine-grained kinematic movements through a language-as-action paradigm. The same rendering applies to real-robot, human-hand and XR-retargeted demos. Second, a parallel dual-training paradigm uses these rationale-augmented trajectories as supervision, training the VLA to predict continuous actions and to articulate the underlying rationale from one shared representation. During deployment, an asymmetric inference strategy strips away the language head, enabling high-frequency control guided by KV-cached expert rationale. Extensive experiments across simulation benchmarks (e.g., LIBERO and VLA-Arena) and real-world manipulation show that StellaVLA effectively exploits heterogeneous demonstrations as structured context, consistently improving both in-distribution performance and OOD robustness, achieving state-of-the-art performance without any robotics pretraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.