Rethinking Vision-Language-Action Model Scaling: Alignment, Mixture, and Regularization
Abstract
Vision–Language–Action (VLA) models show strong promise for generalist robot control. However, it remains unclear whether—and under what conditions—the standard “scale data" recipe translates to robotics, given that robotic training data is inherently heterogeneous across embodiments, sensors, and action spaces. We present a systematic, controlled study of VLA scaling that revisits core training choices for pretraining across diverse robots. Using a representative VLA framework combining a vision–language backbone with flow-matching, we ablate key design decisions under matched conditions in extensive simulation and real-robot experiments. To improve the reliability of real-world evaluation, we introduce a Grouped Blind Ensemble protocol that blinds operators to model identity and randomizes evaluation order within groups. Our analysis targets three dimensions of VLA scaling: (1) Physical alignment: Using SE(3) as a shared action space, we find that end-effector (EEF)-relative actions support positive cross-embodiment transfer. (2) Embodiment mixture: Adding heterogeneous robot datasets can reduce downstream transfer performance, even with the same training budget. (3) Training regularization: We observe that common strategies, such as sensory dropout and multi-stage training, do not consistently improve performance at scale. Together, this study challenges common assumptions about embodied scaling and provides practical, evidence-based guidance for training large-scale VLA policies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.