Discovering Parameterized Motor Skills for Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation, yet their learned behaviors often remain tied to specific task contexts. A key observation is that semantically different manipulation tasks can share recurring short-horizon motor behaviors. Inspired by unsupervised skill discovery in reinforcement learning (RL), we formulate reusable behavior in VLA as discovering, parameterizing, and planning reusable motor skills. We introduce , a framework that discovers recurring local motion patterns from multi-task demonstrations and converts them into detachable parameterized skill modules around a frozen task-adapted VLA. Specifically, a factorized vector-quantized variational autoencoder identifies recurring motion patterns, and each discovered skill is encoded as a low-rank weight update around the frozen supervised fine-tuning (SFT) VLA policy. This parameterization provides compact and detachable skill modules that can be selectively activated while preserving the frozen base policy and its native action interface. During execution, a vision-language router performs skill planning from the current observation and language instruction by choosing either a parameterized skill or the base policy. Experiments on LIBERO and RoboTwin 2.0 show consistent improvements over SFT references, while cross-task skill reuse and real-world evaluation demonstrate that learned skills can transfer across task contexts. Our results point toward a VLA paradigm in which increasingly capable high-level reasoning can leverage reusable motor skills, rather than relearning low-level behaviors for each new task. Code is available at https://anonymous.4open.science/r/0D2F.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.