Policy-Level Recursive Self-Improvement for Embodied AI with a Criticality World Model
Abstract
Despite advances in policy pretraining, embodied AI systems can plateau during task-specific fine-tuning as uniform scenario collection encounters fewer of the remaining failures. We address this problem with a policy-level recursive self-improvement loop: execution outcomes train a criticality world model, whose risk scores guide scenario collection for specialist training. To account for the resulting shift in scenario frequencies, weighted resampling approximately corrects the collection bias. Because specialized updates can weaken nominal behavior, a risk-based gate selects between the specialist and a frozen nominal policy at inference. We extend this procedure to reinforcement learning and behavior cloning. A sampling analysis characterizes the correction, while a controlled study examines the trade-off between critical and nominal coverage. Across five embodied domains, the composed systems reduce failure rates by 59–67% for locomotion and manipulation and 8–25% for VLA benchmarks relative to their respective baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.