Event2Reward: Learning Robotic Rewards from Progress Relations
Abstract
Reward models trained with normalized temporal labels from successful demonstrations can mistake setbacks for progress during exploration, turning losses of progress into positive shaping rewards that hinder policy improvement. Real suboptimal and failed executions offer corrective experience, but their absolute progress is difficult to label. We introduce Event2Reward, a reward model that learns from within-trajectory progress relations supported by task events in real executions. These relations directly supervise the progress predictions used for reward shaping. We combine them with temporal labels from successful demonstrations to anchor the numerical progress scale. We curate an event-annotated dataset from approximately 100K raw real-robot execution records across multiple sources and adapt Event2Reward using a small set of target-task executions spanning successful, suboptimal, and failed outcomes. On 40 tasks across four LIBERO suites, Event2Reward improves suite-average policy success over the strongest baseline in each suite by 9.96–13.76 percentage points under matched per-suite interaction budgets. In real-world experiments with a matched 30-minute wall-clock training budget, Event2Reward achieves an average success rate of 60.0% on three SERL insertion tasks, compared with 37.8% for Robometer, and 90.0% on three DSRL manipulation tasks, compared with 81.7% for Robometer. Ablations support the complementary roles of temporal labels and progress relations in downstream policy learning. We will release our model, dataset annotations, and processing tools.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.