ShareVLA: Share What Is Common, Specialize What Is Needed for Multi-Task Vision-Language-Action Learning
Abstract
Vision-Language-Action (VLA) models adapt pretrained vision-language models to robot control and now solve individual manipulation tasks reliably, yet a single policy that executes many skills remains an open problem. Naive multi-task training suffers from interference among skills, and the prevailing alternative, merging independently trained experts, requires task-specific masks and must leave the deepest layers of the action expert unmerged. This fragility is traced to its source: independently trained action experts diverge from their very first block, so that no component of one expert transfers to another skill and every new skill requires a full training run. We therefore propose ShareVLA, which reverses the order of merging and training. A shared trunk covering of the action expert is trained once on demonstrations from all skills and frozen. Lightweight expert heads are then trained on top of it, one per skill. The shared frozen trunk enables all expert heads to be assembled into a single model with zero merge error by construction, without backbone masks and with only a single backbone forward pass at inference. A lightweight router selects the head. On LIBERO, ShareVLA reaches average success at the same parameter count as the strongest merging-based method while using of its training compute, and exceeds multi-task training of the same architecture by points. These results indicate that multi-skill VLA policies need not merge what was learned apart, and that deciding what to share before specializing is simpler and cheaper.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.