When Learned Rewards Mislead Robotic Reinforcement Learning: Diagnosing and Calibrating Rewards under Distribution Shift
Abstract
Dense rewards provide informative feedback for robotic reinforcement learning (RL), but their reliability can deteriorate once online policy optimization drives the state visitation beyond the reward model’s training distribution. In this work, we first diagnose this failure mode and reveal that online RL induces substantial out-of-distribution (OOD) state visitation, where reward prediction errors are markedly larger than on in-distribution states and can misguide policy optimization. Our key insight is to calibrate rewards according to their estimated reliability, preserving high-confidence predictions while suppressing unreliable ones. We therefore propose ***Confidence-Aware Reward Calibration (CRC)***, which estimates reward confidence from feature-space distances to the reward model’s training data and calibrates low-confidence predictions accordingly. In this way, CRC treats the reward model as a frozen black box and can be seamlessly integrated with different learning-based reward models and RL algorithms without retraining it. Across diverse reward models and RL algorithms, CRC consistently improves online policy learning in both simulation and real-world settings, demonstrating its effectiveness and generality under distribution shift. Code and videos are publicly available at: https://crc.k7m2.workers.dev/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.