acceptodds
Under review as a conference paper at ICLR 2027

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

Abstract

Action decoders in vision-language-action (VLA) models are commonly trained by behavior cloning to reproduce demonstrated action sequences. To enrich this supervision, prior work adds training signals based on future observations and motion. These signals describe executions and outcomes, but do not explicitly identify the local objective an action segment serves under the instruction. We propose Intention Distillation (INDI), which makes this behavior-level intent a supervised intermediate state that action generation is trained to use. A frozen teacher vision-language model (VLM) provides multimodal targets from executed demonstration segments under the task instruction. The student learns to predict these targets at an intermediate decoder layer using only its current observations, instruction, and robot state. Context routing and an intent-mismatch objective encourage downstream predictions to depend on the recovered intent, while visual-outcome and textual-purpose targets provide complementary supervision. The teacher and target-generation modules are removed at deployment. INDI improves GR00T-N1.7 from 64.3% to 84.7% on SimplerEnv-Bridge and from 64.1% to 70.3% on RoboCasa Kitchen, with improvements on across both benchmarks. Across standard scenes, held-out objects, and distractor conditions, average real-world success increases from 62.0% to 68.7%. Representation probes and interventions further show that the recovered states encode task and progress information and influence downstream decoding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.