acceptodds
Under review as a conference paper at ICLR 2027

RE-0: Verified Recursive Improvement of Embodied Code-as-Policy Agents through Local On-Policy Distillation

Abstract

Code-as-Policy agents accomplish long-horizon embodied tasks by generating and executing code, yet continually improving them with teachers that are stronger but not globally reliable remains a key challenge. Existing distillation methods typically treat the teacher's complete behavior as the supervision target and thus misassign training credit on states where the teacher fails. We propose RE-0, a recursively verified policy improvement framework: rather than assuming that the teacher globally outperforms the student, RE-0 requests local corrections from the teacher on the student's own failure histories and checks in the environment whether each correction is genuinely beneficial; verified corrections yield immediate improvement. Building on this, we propose RE-OPD, which turns verified interventions into supervision for on-policy distillation. Only counterfactually verified teacher interventions provide distribution-level supervision, weighted by their measured local benefit, and the improvement they induce is projected back into the standalone student—so both where supervision is applied and how much credit the teacher receives co-evolve with the student policy. We further prove that the student's per-round gain is lower-bounded by its verified intervention gain up to verification and projection error terms. Experiments across multiple Code-as-Policy embodied tasks show that RE-0 improves both teacher-assisted execution and the standalone student, and generalizes to novel robots and scenes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.