acceptodds
Under review as a conference paper at ICLR 2027

Multi-student On-Policy Self-Distillation with Alternating Reinforcement Learning for Robust Vision-Language-Action Models

Abstract

Vision-language-action (VLA) models have emerged as a paradigm for general-purpose robotic control by translating visual observations and natural language instructions into executable actions. Recent reinforcement learning (RL) post-training further improves their task performance through online environment interaction and task-completion rewards. However, VLA policies remain sensitive to prompt variation, exhibiting substantially different success rates even across semantically equivalent instructions. For the same task, high-performing prompts can reliably elicit successful behavior, whereas low-performing prompts often provide few successful trajectories for direct RL. To exploit this performance asymmetry, we propose Multi-student On-policy Self-distillation with Alternating Reinforcement learning (MOSAR). MOSAR uses a high-performing prompt as a teacher for multiple semantically equivalent student prompts and distills teacher actions at student-visited observations, enabling cross-prompt behavior transfer within a shared VLA policy. For multi-task training, MOSAR adaptively weights distillation according to task-specific teacher–student performance gaps and alternates RL and distillation with separate optimizer states, allowing policy improvement and cross-prompt knowledge transfer to reinforce each other throughout training. From extensive evaluations on RoboLab, REALM, and a real Franka robot, MOSAR consistently improves success on unseen prompts over RL- and OPSD-based baselines while maintaining strong performance on targeted and unseen tasks. Additional evaluations under visual distractors further demonstrate robustness beyond the prompt variations explicitly addressed during training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.