Stabilizing RL Fine-Tuning of Diffusion VLAs with Sparse MoE Action Heads: Advantage- and Noise- Aware Load Balancing and Geometric-Mean Ratio
Abstract
Sparse mixture-of-experts (MoE) is an effective architecture for scaling vision- language-action (VLA) policies, but imitation learning alone is limited by demon- stration coverage and cannot fully exploit MoE capacity under distribution shift. Reinforcement learning (RL) fine-tuning can leverage task feedback, yet directly applying RL to diffusion MoE-VLAs is unstable because continuous denoising dynamics interact with discrete expert routing. We diagnose two failure modes: (1) advantage-signed routing updates sharpen experts for positive-advantage sam- ples and flatten them for negative-advantage samples, while our measurements show larger routing changes at late low-noise steps; and (2) expert switching along the denoising chain causes step-wise likelihood-ratio deviations to accumulate, making the resulting trajectory ratios heavy-tailed and frequently clipped. To ad- dress these issues, we propose two minimal interventions. First, as a loss term, Advantage- and Noise-Aware Load Balancing (ANA-LB) applies noise-scaled load balancing only to trajectories with negative group-relative advantage, sta- bilizing relatively worse samples while preserving positive-advantage expert spe- cialization. Second, for ratio computation, Geometric-Mean Ratio (GMR) uses a geometric-mean surrogate for the accumulated step-wise log-ratios, trading a small amount of exact importance-weighting fidelity for reduced ratio variance. Experiments on LIBERO, CALVIN, RoboTwin 2.0, and real-robot tasks show competitive overall performance, with the clearest gains in several long-horizon settings, while reducing routing collapse, excessive ratio clipping, and training oscillation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.