AdaMoE-VLA: Adaptive Mixture of Action Experts for Vision-Language-Action Learning
Abstract
Vision-Language-Action (VLA) models show strong potential as generalist robot policies, yet enhancing their action-generation capability remains challenging due to the diversity of robot behaviors and the need for fine-grained action modeling. While Mixture-of-Experts (MoE) offers an efficient way to increase model capacity through sparse activation, its role in VLA is not simply generic model scaling. Long-horizon manipulation trajectories often contain heterogeneous and frequently switching action modes, such as reaching, grasping, aligning, inserting, and recovery behaviors, making multi-expert action modeling especially suitable for VLA policies. However, directly applying conventional MoE to VLA action heads introduces an optimization conflict: the router is required to both select experts under a load-balancing objective and provide precise state-dependent expert weighting for continuous action generation. This coupling can interfere with fine-grained action learning, particularly when behavior modes change frequently. Therefore, we propose AdaMoE-VLA, an adaptive mixture of action experts for VLA learning. AdaMoE-VLA decouples expert selection from expert weighting by preserving the standard router for trainable sparse selection while introducing a lightweight Scale Adapter for task-conditioned expert reweighting. This design mitigates gradient interference between load balancing and action prediction, enabling task-adaptive expert reweighting while preserving real-time inference. Across simulation and real-world benchmarks, AdaMoE-VLA consistently improves -series VLA baselines, and its gains become more pronounced as action-mode transitions become more frequent. It reaches 98.3% on the LIBERO benchmark, while delivering larger improvements in more behaviorally diverse settings, with a 9.3% gain on RoboTwin 2.0 and a 20% – 21.5% absolute gain in real-world experiments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.