Self-Improving Robot Skills for Long-Horizon Tasks via Recursive State-Aware Credit Assignment
Abstract
Long-horizon robot programs can fail when early operations leave states that compromise later execution. We present Recursive State-Aware Credit (RSAC) for improving robot skills through program revision, with the language model and action primitives fixed. Continuations of unchanged program suffixes from saved stage exits estimate how intermediate states support complete-task success. State-aware credit (SA) separates local passage deficits from poor exit-state quality. RSAC recursively compares gains in terminal-outcome dependence across stage boundaries to select an editing region for the current program. Execution experience guides localized revisions, and paired full-task validation determines adoption for the next iteration. Across eight simulated quadrupedal robot tasks, with ten independent optimization runs per task and configuration, RSAC achieves 44.4% task-averaged success, compared with 30.3% for ASPIRE and 35.4% for one-step localization. Controlled interventions and online revisions connect the diagnostic evidence to concrete program changes and improved task performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.