dVLA-RL: Reinforcement Learning over Denoising Trajectories for Discrete Diffusion Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models have established a powerful paradigm for generalist robotic manipulation by grounding control into the semantic reasoning of large Vision-Language Models (VLMs). Prevailing architectures typically model actions continuously via diffusion or flow processes, or discretely through either autoregressive generation or parallel decoding. Recently, Discrete Diffusion VLAs (dVLAs) have emerged as a distinct alternative, unifying vision, language, and action into a single discrete token space via masked generative modeling. While this paradigm successfully combines iterative action refinement with unified representations, its training has thus far been restricted to Supervised Fine-Tuning (SFT), leaving the potential of Reinforcement Learning (RL) for further policy refinement largely unexplored. A fundamental challenge in RL for dVLAs is that the marginal probability of the final action remains intractable. To address this challenge, we propose dVLA-RL, which shifts policy optimization from the intractable marginal likelihood of the final action to a tractable trajectory-level objective defined over the sampled denoising path. By factorizing this objective across denoising transitions, dVLA-RL naturally accommodates variable denoising horizons. Leveraging this flexibility, we introduce a unified step scheduling approach for complex multi-task learning, tailoring denoising steps to specific task complexities to maximize both success rates and computational efficiency. Extensive evaluations show that dVLA-RL achieves 99.7% success on LIBERO and establishes strong VLA-based results on RoboTwin 2.0, improving over the SFT baseline by 30.6 percentage points while remaining competitive with strong World-Action Model baselines. On real-world tasks, dVLA-RL improves the aggregate score by 25.8 percentage points over the MM-ACT baseline, with performance comparable to the representative continuous-action VLA .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.