Atomic-to-Composite Distillation via Task-Wise Progress Potentials and Hidden-State Alignment in Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models achieve strong performance on individual manipulation skills, yet their success drops substantially when these skills must be executed sequentially as composite tasks. We identify two key challenges underlying this atomic-to-composite gap. First, composite behaviors are difficult to learn: collecting long sequence demonstrations is costly, while reinforcement learning (RL) with sparse binary rewards provides little learning signal for trajectories that achieve only partial progress. Second, composite instructions do not explicitly indicate which atomic task should be active at each stage, creating ambiguity at subtask transitions. Moreover, the same skill succeeds far less often under a composite instruction than under its own atomic instruction, suggesting that composite conditioning fails to reliably invoke competence the policy already possesses. We propose ATLAS, an end-to-end RL framework that leverages the policy's own atomic-task competence to address both challenges. To address sparse credit assignment, task-wise progress potential shaping learns reusable progress potentials from successful and failed atomic rollouts and composes them across stages to provide dense, phase-aware rewards while preserving the optimal policy. To bridge the representation gap, phase-conditioned atomic-to-composite hidden-state alignment aligns composite-conditioned representations with their corresponding atomic-conditioned representations at each stage, allowing the composite-conditioned policy to directly exploit skills already encoded by the atomic policy without inference-time planners or prompt switching. These complementary mechanisms bridge atomic and composite learning within a unified training framework. Across simulation and real-robot experiments, ATLAS consistently outperforms zero-shot VLA baselines, standard PPO, and existing progress-reward approaches, achieving gains of up to 33pp on RoboLab and 40pp in real-world settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.