acceptodds
Under review as a conference paper at ICLR 2027

Auditing Learned Robot Rewards

Abstract

Learned visual rewards can score near-goal failures highly, allowing PPO to increase proxy reward while task success falls. Fixed-P8 runs show a 9.9 pp ManiSkill peak-to-final success decline and a matched-proxy residual that turns negative. Label-Selective Audit and Refresh (LSAR) uses random gold audits of top-scored trajectories and refreshes the reward when their precision falls below a reference level. At a 1,000-label cap, LSAR ends at 72.3% on ManiSkill and 40.8% on LIBERO-Long, versus 66.9% and 33.4% for periodic top-score labeling, using a mean of 765 and 637 method-requested gold labels, respectively; shared setup and diagnostic labels are excluded. Shadow precision, held-out audits, matched-cost controls, and component analyses characterize the refresh mechanism and its label cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.