Stochastic Self-Distillation for Vision-Language-Action Models
Abstract
Flow-based vision–language–action (VLA) models achieve strong performance in robotics through imitation learning, yet fine-tuning them with on-policy reinforcement learning (RL) is limited. In group-based RL methods like Group Relative Policy Optimization (GRPO), every rollout step requires multi-step ODE integration, and the sampler's limited stochasticity leads to rollout groups with low variance. We propose VLA-SSD, which addresses both high inference latency and limited stochasticity. In particular, our method self-distills a pretrained flow-matching action head into a stochastic one-step policy by conditioning a single-jump generator on a Brownian path. We derive a conditional action density via our stochastic one-step policy, enabling efficient post-training via likelihood-ratio RL algorithms. We show that distilling VLA-SSD from the reduces the action head's forward-pass time by times, with a relative decrease in LIBERO success rate. Leveraging the increased sample diversity of VLA-SSD during group generation, we show that post-training with GRPO from a few-shot checkpoint raises LIBERO success from 77.1% to 93.0%, achieving a relative improvement over four-step GRPO.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.