MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation
Abstract
Vision-Language-Action (VLA) models trained on large-scale robotics datasets promise robust control across diverse domains and embodiments, yet they frequently fail zero-shot in novel scenarios. We introduce Mixture-of-Skills VLA (MoS-VLA), a framework that formulates robot manipulation policies as continuous linear combinations of learned basis functions. Pretrained jointly on the Open X-Embodiment data mixture, MoS-VLA constructs a skill space for rapid adaptation. At test time, adaptation in our evaluated robot settings uses a single expert demonstration. We infer the optimal mixture coefficients via a lightweight, gradient-free convex optimization that minimizes L1 action error. MoS-VLA achieves 70-100% success in simulated and real-robot tasks, where baseline success is 0% in all but one comparison. It also achieves lower action-prediction error than two pretrained VLAs and two in-context baselines on their evaluated datasets, while adapting at a fraction of the compute of gradient-based finetuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.