DemoPrior: Learning Unseen Robot Tasks from a Single Demonstration with Action-Space Priors
Abstract
We propose DemoPrior, a one-shot video-to-skill framework that enables a learned vision-language-action (VLA) policy to acquire unseen manipulation skills from a single demonstration video. Existing approaches either learn adaptation mechanisms whose transfer remains tied to the tasks and variations covered during training, or use a test-time demonstration mainly as a weak conditioning signal or trajectory initialization. We instead propose to learn a reusable demo-to-skill transformation that converts information extracted from a new demonstration into executable behavior. Our key insight is that a demonstration provides two complementary forms of task information: semantic structure specifying what entities and interactions matter, and task-relative spatial structure specifying how the behavior should unfold in space. DemoPrior extracts these semantic and spatial cues from the video, transfers them to the current scene, and maps them into an action-space skill prior. A flow-based VLA is trained to refine this prior using closed-loop observations, learning how to turn video-derived structure into robot actions rather than generating a new skill from scratch. At test time, a single demonstration can therefore introduce a new skill without any task-specific policy update. Across our unseen-task evaluations, DemoPrior consistently improves demonstration-based skill transfer, and achieves a 46.7% success rate on real-world unseen tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.