acceptodds
Under review as a conference paper at ICLR 2027

STEPS: Scene-Grounded Task Escalation through Progressive Skill Learning toward Open-Ended Goals

Abstract

Code-as-Policy robot agents increasingly improve through practice, yet evalua- tion on predefined tasks makes it difficult to distinguish genuine capability acqui- sition from adaptation to practiced instances. We present STEPS, a framework for evaluating self-improvement on human-specified goals such as “tidy the office desk”. STEPS generates physically valid, oracle-solvable scenes; defines the full goal as an L5 task and derives L1-L4 diagnostic tasks probing subtasks, compo- nent capabilities, and constraints; and evaluates a frozen agent on the practiced L5 task, pose-perturbed L5 instances, and held-out lower-level tasks. Across five goals and 105 tasks, a skill library developed with human assistance achieves over 90% success on practiced L5 tasks but performs substantially worse on regener- ated instances and held-out L3-L4 tasks. On one evolution route, skills distilled from L1-L4 practice usually pass isolated validation but are rarely invoked when solving the full goal. Despite this fitting, some learned capabilities transfer be- yond the development scenes to external manipulation tasks, and weaknesses ex- posed by STEPS provide actionable targets for capability updates that further im- prove cross-domain performance. These results show that practiced-task success alone can overstate capability acquisition and should be complemented with goal- derived diagnostics that reveal what transfers, where improvement breaks down, and what capability to target next.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.