acceptodds
Under review as a conference paper at ICLR 2027

GRVM: Goal-Relative Value Modeling Beyond Temporal Progress for Robot Manipulation

Abstract

Robot value models increasingly provide feedback for execution evaluation and policy improvement. However, supervision based on task advancement does not necessarily reflect whether the current physical state has become more favorable for task completion. This distinction is particularly important around failures and recoveries, where temporal progress and physical state quality can diverge. We introduce Goal-Relative Value Model (GRVM), which separately predicts task progress, goal-relative state quality, and trajectory outcome from visual observations and language. A human-audited annotation pipeline constructs dense goal-relative supervision using end-effector poses at successful stage completion as local goals and extends this supervision to failed and recovered executions through failure-aware relabeling. We pretrain GRVM on 286,316 annotated trajectories (2,612 hours), including 17,286 failure/recovery trajectories. We also construct a balanced, multi-embodiment benchmark of 1,000 held-out trajectories to evaluate continuous-label agreement and local state ordering around failures and recoveries. GRVM improves triplet accuracy by 58.0% and recovery-direction accuracy by 39.7% relative to the strongest competing baseline for each metric. With GRVM kept frozen, its feedback increases mean success across four LIBERO suites from 89.8% to 94.3% over sparse-reward RL while accelerating learning. GRVM-guided online RL raises mean success from 55.0% (SFT) to 95.0% across four real-robot tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.