ReLoop: Progress Supervision from LLM-Generated Reward Program Execution
Abstract
Large language models (LLMs) can produce executable reward programs from task descriptions, but the downstream reward interface typically exposes only each program's scalar return. Distances and contact predicates computed inside the program may reveal differences in task progress even when scalar rewards coincide. We introduce ReLoop, which records these runtime values and maps them to a progress potential and trace state. Potential changes supply a training signal, while a frontier bonus and memory of observed success use past traces to provide additional signals. ReLoop leaves the generated reward logic and policy observations unchanged and requires no further LLM calls. Across eight manipulation tasks, it improves mean final success by 17.1 percentage points over direct optimization of the same programs and by 7.0 points over a control matched for reward scale and success feedback. A control built from the scalar return favors the structured interface on average. On three tasks, the potential term retains the endpoint advantage over the matched control without the two memory terms. ReLoop thus turns intermediate values from a generated reward program into progress supervision during policy training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.