Auditing Learned Robot Rewards
Abstract
Learned visual rewards can score near-goal failures highly, allowing PPO to increase proxy reward while task success falls. Fixed-P8 runs show a 9.9 pp ManiSkill peak-to-final success decline and a matched-proxy residual that turns negative. Label-Selective Audit and Refresh (LSAR) uses random gold audits of top-scored trajectories and refreshes the reward when their precision falls below a reference level. At a 1,000-label cap, LSAR ends at 72.3% on ManiSkill and 40.8% on LIBERO-Long, versus 66.9% and 33.4% for periodic top-score labeling, using a mean of 765 and 637 method-requested gold labels, respectively; shared setup and diagnostic labels are excluded. Shadow precision, held-out audits, matched-cost controls, and component analyses characterize the refresh mechanism and its label cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.