acceptodds
Under review as a conference paper at ICLR 2027

OneDemo-α: From Spatial Reasoning to One-Shot Robot Skill Transfer

Abstract

Imitation learning and pretrained robot models pursue generalization through large-scale training, yet rapidly transferring a newly demonstrated skill across objects and goals remains challenging. Vision-language models (VLMs) with strong spatial reasoning offer a complementary route to adaptation, but repeated inference and action generation during execution can incur substantial latency. Our key insight is to combine reusable motion experience from a demonstration with a spatially proficient VLM’s ability to determine how that motion should change. We introduce OneDemo-α, a framework for one-shot skill acquisition and transfer without task-specific training. From one RGB-D demonstration with synchronized robot poses, the VLM predicts an editable skill-adaptation vector α for the target scene and goal. Its functional bindings, motion assignments, and sparse edits guide a geometric executor, preserving reusable motion while concentrating inference on transfer decisions. We further introduce OneDemo-Bench, which systematically varies object position, appearance, size, morphology, and task requirements to evaluate transfer from a fixed demonstration. Experiments show that OneDemo-α achieves higher transfer success than trajectory-transfer baselines and substantially shorter execution times than direct VLM control, enabling efficient one-shot skill reuse without task-specific training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.