PCMP: Learning Pace-Controllable Motion Priors for Robotic Manipulation
Abstract
Vision-language-action (VLA) models connect visual–language understanding to robotic manipulation. However, motion-level generalization is limited in multi-task settings, and learning controllable execution pace from a single dataset remains difficult. These limitations motivate a reusable motion backbone that preserves task structure while responding to user-specified execution rates. We introduce Pace-Controllable Motion Priors (PCMP), a hierarchical framework that separates motion content, current progress, and desired future progression. PCMP encodes semantically aligned demonstrations into reconstructable motion latents and learns a progress-conditioned flow-matching prior for action generation. Complementary spatial and progress-schedule augmentations vary motion geometry and timing independently. A visual-language upper model predicts the motion latent and current progress, connecting the prior to closed-loop manipulation while retaining an explicit interface for pace commands. Experiments on RoboTwin show that the learned prior provides an effective motor backbone and retains useful execution capabilities when coupled to visual perception. Motion-latent ablations support the contribution of the trajectory representation, while speed-control experiments reveal a systematic response to user-specified rates: execution accelerates as the commanded rate increases. These findings support learning reusable motion structure separately from visual grounding and demonstrate a promising route toward manipulation with user-controlled execution pace.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.