PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation
Abstract
Vision-Language-Action (VLA) models offer a promising route to general-purpose robotic policies, yet downstream adaptation remains data-intensive, with limited transfer to new task compositions. Direct instruction-to-control fine-tuning treats multi-stage demonstrations as complete task trajectories, offering little structure for reusing recurring interaction patterns. We propose PrimitiveVLA, a primitive-centric Disassembly-and-Assembly framework. During fine-tuning, a vision-language model (VLM) infers primitive sequences, while a large language model (LLM) localizes their temporal boundaries to construct primitive-aligned samples. During inference, the VLM generates a primitive flow; the VLA executes it in a closed loop with an LLM-generated causal switching program. Unified Multimodal Representation (UMR) maps instance-specific descriptions to unified primitive instructions while grounding objects and targets through mask-augmented observations. Across Libero, RLBench, and real-robot experiments, PrimitiveVLA improves data efficiency and zero-shot transfer to unseen task compositions and long-horizon tasks while maintaining or improving in-distribution performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.