acceptodds
Under review as a conference paper at ICLR 2027

ROAD-VLA: Robust Online Adaptation via Self-Distillation for Vision-Language-Action Models

Abstract

Online reinforcement learning can adapt vision-language-action (VLA) models to new deployment conditions, but scalar advantages provide coarse supervision for their high-dimensional, structured action predictions. We propose ROAD-VLA, an advantage-guided self-distillation framework that turns rollout feedback into a structured teacher directly in the policy's native action space. After each on-policy rollout, ROAD-VLA takes a proximal mirror-descent step from the rollout prediction, moving toward better-than-expected executed behavior and away from worse-than-expected behavior. To improve reliability under distribution shift, a frozen reference critic is incorporated only when its advantage estimate agrees in sign with the online critic. This yields a unified construction for discrete action tokens, continuous action chunks, and flow-matching velocity predictions, without requiring an external expert. We derive a conditional policy-improvement bound under explicit calibration and stability assumptions, providing a common theoretical interpretation across all three action representations. Across ManiSkill environments spanning visual, compositional, and execution shifts, ROAD-VLA improves OpenVLA-7B and OpenVLA-OFT over PPO and outperforms language-guided and external-teacher distillation baselines. We further demonstrate the same mechanism with the flow-matching pi_0.5 policy on LIBERO. Across these settings, ROAD-VLA enables faster online adaptation while maintaining higher policy entropy and lower late-stage variance than PPO.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.