Vision-Language Judges Prevent Reward Hacking in Robot Manipulation RL
Abstract
Reinforcement learning (RL) is a major part of the training pipeline of the most powerful models today. However, designing good RL environments is hard, and models often find creative ways to achieve high reward without fulfilling the designer's intent, known as reward hacking. This poses a major risk to the safety of AI systems. A natural but little-studied countermeasure is to let a trusted model judge each rollout during training and intervene, e.g., by penalizing or discarding flagged rollouts. We introduce three manipulation environments in which Vision-Language-Action models (VLAs) learn to reward hack during RL, and test such interventions with Vision-Language Models (VLMs) as judges. First, using oracle judges that read the true environment state, we find that small penalties suppress hacking nearly as well as strong ones while costing far less task performance when the judge falsely flags honest rollouts. We then build judges from VLMs spanning small to frontier models and improve them by letting them output scores instead of binary verdicts and giving them context and tools. Finally, we match the oversight performance of the strongest VLM judges at over lower cost by training small task-specific judges, open-weight VLMs and probes, on the frontier judge's verdicts. Beyond our setting, VLMs overseeing VLAs is a case study of a judge supervising a model that differs from it in training and capabilities, and a starting point for building such oversight more broadly.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.