LADDER: Language and Affordance Driven Demonstration Engine for Robots
Abstract
Robot learning has moved closer to generalist agents that can act in everyday environments, but efficiently collecting high-quality manipulation data for training remains a challenge. Real-world manipulation data collection often relies on human demonstrations, including robot teleoperation and handheld devices that record human manipulation trajectories. While data collection in simulation mitigates this labor-intensive burden, existing approaches often rely on manually acquired affordances for manipulation or carefully designed reward shaping for reinforcement learning before data collection. To efficiently collect manipulation data, we present LADDER, an automated pipeline for demonstration generation given a natural language instruction. In this pipeline, a manipulation knowledge base is built using parameterized instances on which affordances are formulated to support automatic task execution; a multi-agent workflow decomposes a natural language instruction into verified atomic tasks and automatically constructs a scene; affordances of each instance are then mapped to a corresponding asset and extended to a sequence of key poses; trajectories connecting the key poses are curated and used to guide robots to perform each atomic task. After a successful execution, the trajectory and related data are generated. LADDER collects demonstration data on 100 tasks of RoboTwin 2.0 and RoboGen, demonstrating that it improves generation efficiency for both rigid-body and articulated tasks. Furthermore, Vision-Language-Action (VLA) models fine-tuned on datasets synthesized by LADDER exhibit high performance, validating the high quality of our collected demonstrations. Our website is available at https: //anonymous.4open.science/w/ladder-project-D8E2/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.