acceptodds
Under review as a conference paper at ICLR 2027

CriticHack: Evaluating Visual Rewards Under Robot Policy Optimization

Abstract

Learned visual rewards guide robot policy optimization, but can assign high scores to executions that act on the wrong object without completing the requested task. We present CriticHack, a closed-loop study of visual reward optimization in simulated drawer manipulation and identity-sensitive stacking. Our central finding is that optimizing a learned visual reward can increase wrong-object failure rates even as reward and task success improve. In a supervised-initialization study, three full-denoiser fine-tuning runs of a diffusion policy with no prior reward exposure, optimized against Robometer on a drawer task, raise simulator-measured success and wrong-object failure rates by 9.6 and 11.8 percentage points (95% CIs [6.2, 13.0] and [8.1, 15.6]) on 512 evaluation seeds, alongside higher reward. A warm-start study shows the same pattern from a policy previously optimized against learned rewards: across five runs, success and wrong-object failures rise by 8.4 and 3.8 points on average. Constrained-policy experiments, which optimize only a selector over fixed diffusion-noise sequences, reproduce reward gain and wrong-object amplification across two task–reward pairings and two optimizers. Reward-ranking analysis further shows that strong aggregate success–failure discrimination can conceal semantic misranking. On collected trajectories, Robometer assigns higher average scores to wrong-object failures than to successes; strong discrimination against other failures masks this scoring error. As a diagnostic intervention, reweighting the critic reward with a frozen verifier trained on task-specific outcome labels improves task completion and reduces wrong-object failures relative to critic-only training under the fine-tuning sampler. Verifier-only controls identify task-specific outcome supervision as a substantial source of recovery. These results show that gains in reward and task success can conceal the amplification of specific semantic failures.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.