Multi-student On-Policy Self-Distillation with Alternating Reinforcement Learning for Robust Vision-Language-Action Models
Abstract
Vision-language-action (VLA) models have emerged as a paradigm for general-purpose robotic control by translating visual observations and natural language instructions into executable actions. Recent reinforcement learning (RL) post-training further improves their task performance through online environment interaction and task-completion rewards. However, VLA policies remain sensitive to prompt variation, exhibiting substantially different success rates even across semantically equivalent instructions. For the same task, high-performing prompts can reliably elicit successful behavior, whereas low-performing prompts often provide few successful trajectories for direct RL. To exploit this performance asymmetry, we propose Multi-student On-policy Self-distillation with Alternating Reinforcement learning (MOSAR). MOSAR uses a high-performing prompt as a teacher for multiple semantically equivalent student prompts and distills teacher actions at student-visited observations, enabling cross-prompt behavior transfer within a shared VLA policy. For multi-task training, MOSAR adaptively weights distillation according to task-specific teacher–student performance gaps and alternates RL and distillation with separate optimizer states, allowing policy improvement and cross-prompt knowledge transfer to reinforce each other throughout training. From extensive evaluations on RoboLab, REALM, and a real Franka robot, MOSAR consistently improves success on unseen prompts over RL- and OPSD-based baselines while maintaining strong performance on targeted and unseen tasks. Additional evaluations under visual distractors further demonstrate robustness beyond the prompt variations explicitly addressed during training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.