Chain-of-Primitives Shapes Robot-Agent Adaptation
Abstract
Vision-language models (VLMs) increasingly control robots through manipulation primitives, ranging from low-level commands to geometric targets and semantic skills. How these capabilities are exposed to the agent determines which parts of spatial reasoning and control the model must solve itself. Existing systems, however, typically fix this interface by design. We introduce PrimitiveSuite, which exposes a common set of manipulation capabilities through three agent-facing interfaces: direct control, geometric specification, and semantic invocation. We find that no single interface is consistently best: the preferred abstraction varies across models, tasks, and even manipulation phases. Rather than prescribing this choice, we allow agents to compose heterogeneous primitives during execution and represent their decisions as a Chain-of-Primitives (CoP). Successful CoPs provide a natural form of in-context experience, demonstrating how available primitives should be selected, parameterized, and sequenced. CoP-based ICL substantially improves composition, while visual demonstrations alone provide little benefit. GPT-6 Astra reaches 82.4% success on LIBERO-Pro, and Claude Opus 5 reaches 92.3% on RoboMME, outperforming representative VLA, coding-agent, and VLM-harness baselines. Our results shift the design question from choosing a fixed robot interface to enabling agents to adaptively compose interfaces from experience.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.