acceptodds
Under review as a conference paper at ICLR 2027

When Learned Rewards Mislead Robotic Reinforcement Learning: Diagnosing and Calibrating Rewards under Distribution Shift

Abstract

Dense rewards provide informative feedback for robotic reinforcement learning (RL), but their reliability can deteriorate once online policy optimization drives the state visitation beyond the reward model’s training distribution. In this work, we first diagnose this failure mode and reveal that online RL induces substantial out-of-distribution (OOD) state visitation, where reward prediction errors are markedly larger than on in-distribution states and can misguide policy optimization. Our key insight is to calibrate rewards according to their estimated reliability, preserving high-confidence predictions while suppressing unreliable ones. We therefore propose ***Confidence-Aware Reward Calibration (CRC)***, which estimates reward confidence from feature-space distances to the reward model’s training data and calibrates low-confidence predictions accordingly. In this way, CRC treats the reward model as a frozen black box and can be seamlessly integrated with different learning-based reward models and RL algorithms without retraining it. Across diverse reward models and RL algorithms, CRC consistently improves online policy learning in both simulation and real-world settings, demonstrating its effectiveness and generality under distribution shift. Code and videos are publicly available at: https://crc.k7m2.workers.dev/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.