Vision-Language Reward Models Confuse Near-Misses with Successes in Robot Manipulation
Abstract
Vision–language models (VLMs) are increasingly used as reward models for robot manipulation, where they filter and rerank a policy's rollouts. They are typically evaluated on how well they separate successful rollouts from failed ones. Most failures in these evaluations are obvious, such as a dropped object, yet reward models are now often applied to policies that are already competent, where nearly every rollout ends close to success. To measure this setting, we find for each task in simulation how far the object can be moved before the task no longer counts as complete, and we create failures just past that limit (near-misses) and far beyond it (gross failures). We test nine judges on four LIBERO suites. Every judge is worse at telling successes from near-misses than from gross failures. This also holds for failures that a robot policy (OpenVLA) makes on its own, not only for the failures we create. When a judge picks the best of three rollouts, near-misses erase on average 24–62% of its advantage over picking at random, depending on the suite. We also create rollouts in which the object is moved but the task still succeeds. Judges score these lower than untouched successes, which suggests that they track how far the object moved rather than whether the task was done. Across the judges we test, near-miss accuracy varies more from task to task than from judge to judge, so a benchmark's score depends heavily on which tasks it includes. Fine-tuning on our near-misses helps on new tasks from the suites it was trained on but not on an unseen suite. We recommend that reward-model benchmarks report near-miss accuracy separately and show how it varies with task difficulty.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.