acceptodds
Under review as a conference paper at ICLR 2027

PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation

Abstract

Vision-Language-Action (VLA) models offer a promising route to general-purpose robotic policies, yet downstream adaptation remains data-intensive, with limited transfer to new task compositions. Direct instruction-to-control fine-tuning treats multi-stage demonstrations as complete task trajectories, offering little structure for reusing recurring interaction patterns. We propose PrimitiveVLA, a primitive-centric Disassembly-and-Assembly framework. During fine-tuning, a vision-language model (VLM) infers primitive sequences, while a large language model (LLM) localizes their temporal boundaries to construct primitive-aligned samples. During inference, the VLM generates a primitive flow; the VLA executes it in a closed loop with an LLM-generated causal switching program. Unified Multimodal Representation (UMR) maps instance-specific descriptions to unified primitive instructions while grounding objects and targets through mask-augmented observations. Across Libero, RLBench, and real-robot experiments, PrimitiveVLA improves data efficiency and zero-shot transfer to unseen task compositions and long-horizon tasks while maintaining or improving in-distribution performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.