BABYSTEPS: Learning Task Meaning Along The Way
Abstract
A robot can change its behavior after failure while retaining a mistaken inter- pretation of the demonstrated task. We study this initial-belief anchoring and introduce BABYSTEPS, which trains a vision-language supervisor to revise task interpretations from demonstrations and execution history while guiding a frozen vision-language-action executor. During offline training, a teacher uses the canon- ical task instruction to write a reflection on the demonstration, initial instruc- tion, and failed rollout. We pair this reflection with the canonical instruction for supervised fine-tuning on student-visible histories, without requiring successful correction rollouts. At deployment, the supervisor generates reflections and re- vised instructions without teacher access. Our Qwen3.5-9B supervisor achieves 77.5% Success@5 on LIBERO tasks held out from supervisor training, exceeding Zero-shot and Instruction-only SFT by 30 and 10 percentage points, respectively. Under controlled incorrect-instruction interventions, both BABYSTEPS and a neutral-reflection control eliminate repairs that express only the implanted task, but BABYSTEPS more often recovers the true task without implanted compo- nents: 20/60 repairs versus 8/60. These results distinguish abandoning an incor- rect interpretation from revising toward the demonstrated task
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.