acceptodds
Under review as a conference paper at ICLR 2027

Learning Optimal Robot Rewards from Imperfect Trajectories with Quasimetrics

Abstract

Video reward models for robot learning are typically trained to predict temporal progress, the relative position of each frame within a demonstration. However, this target measures the demonstrator’s value rather than the optimal one. Pauses and detours therefore distort the target, and failed attempts yield no target at all, since the time to success is never observed. We propose Quasimetric Reward Model (QRM), which targets the optimal goal-reaching value by learning from bounds rather than regression targets. Each observed video segment is a feasible path whose return provides a lower bound on the optimal value, even when the trajectory is inefficient or failed. A quasimetric model combines these bounds through the triangle inequality to tighten its value estimates. Such a value is defined for a single goal state, whereas a language instruction admits many successful outcomes and provides no goal image for a new task. QRM therefore distills the value of reaching the nearest successful training outcome into a compact model conditioned on visual history and language. Its dense reward complements the sparse success signal during reinforcement learning. On MetaWorld tasks unseen during reward model training, QRM improves its policy success rate with imperfect videos that degrade temporal-progress models, outperforming vision-language reward models 50× its size. For fine-tuning a pretrained robot policy, QRM yields 1.8× the success rate of the strongest baseline on held-out LIBERO-90 tasks and matches or exceeds every baseline on real-robot manipulation tasks learned from few demonstrations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.