acceptodds
Under review as a conference paper at ICLR 2027

Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs

Abstract

Vision-language-action (VLA) models often struggle to generalize to skill combinations absent from their fine-tuning demonstrations, even when every constituent skill has been demonstrated. We focus on a vision shortcut as one failure mode: during fine-tuning, visual observations can serve as a proxy for the instruction, so a policy may execute a demonstrated combination associated with similar observations rather than the instructed combination. This motivates counterfactual pairs that hold a demonstration observation fixed while changing the instruction to specify an undemonstrated combination. These pairs lack demonstrated actions for supervision. Crucially, the currently required skill has already been demonstrated, but actions from those executions cannot be transferred directly because the same skill can require different actions across observations. We propose CRAFT, which transfers supervision from demonstrated executions of the required skill to counterfactual pairs through reusable skill representations. Across three VLA models and two simulation benchmarks, CRAFT improves success on undemonstrated combinations while maintaining high success on demonstrated ones. CRAFT also improves compositional generalization on a real robot.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.