acceptodds
Under review as a conference paper at ICLR 2027

ReLoop: Progress Supervision from LLM-Generated Reward Program Execution

Abstract

Large language models (LLMs) can produce executable reward programs from task descriptions, but the downstream reward interface typically exposes only each program's scalar return. Distances and contact predicates computed inside the program may reveal differences in task progress even when scalar rewards coincide. We introduce ReLoop, which records these runtime values and maps them to a progress potential and trace state. Potential changes supply a training signal, while a frontier bonus and memory of observed success use past traces to provide additional signals. ReLoop leaves the generated reward logic and policy observations unchanged and requires no further LLM calls. Across eight manipulation tasks, it improves mean final success by 17.1 percentage points over direct optimization of the same programs and by 7.0 points over a control matched for reward scale and success feedback. A control built from the scalar return favors the structured interface on average. On three tasks, the potential term retains the endpoint advantage over the matched control without the two memory terms. ReLoop thus turns intermediate values from a generated reward program into progress supervision during policy training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.