acceptodds
Under review as a conference paper at ICLR 2027

Dynamic Episode Alignment: Sample-Efficient Reinforcement Learning of Vision-Language-Action Models

Abstract

Online reinforcement learning (RL) can improve vision-language-action (VLA) models beyond imitation learning, but many methods rely on a final success-or-failure reward that does not reveal which actions helped. RL can therefore reinforce hesitations in successes and discourage productive actions in failures, requiring many task attempts to resolve this conflicting credit. We propose Dynamic Episode Alignment (DEAL), which uses the policy's own successes to identify productive and wasted steps without an extra scoring model. DEAL uses dynamic programming (DP) to align each new episode with the policy's shortest success in the same scene, matching steps by the policy's own representations. RL then updates only productive steps in successes and wasted steps in failures, so successes no longer reinforce hesitations and failures no longer discourage productive actions. In a tabular model under simplifying assumptions, we prove that selecting steps reduces the episodes needed to reliably reinforce productive actions from linear to constant in the number of subgoals. To support DEAL's updates, we introduce a recurrent Gaussian action head that provides exact action likelihoods. Combining DEAL with PPO and the recurrent Gaussian action head yields 97.6% success on LIBERO with one demonstration per task and 80.4% on RoboTwin 2.0, versus SimpleVLA-RL's 96.9% and 66.4%. From the same starting policies on three RoboTwin 2.0 tasks, DEAL improves PPO's sample efficiency and raises mean success by a relative 26% after equal training. On a humanoid with two dexterous hands, online RL with DEAL exceeds EgoVLA's success by a relative 46.8% on unseen backgrounds.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.