acceptodds
Under review as a conference paper at ICLR 2027

Calibrate Before You Learn: Coverage-Aware Rewards for Offline RL

Abstract

Offline reinforcement learning has access to trajectory outcomes before training begins, but early temporal-difference updates still rely on a critic that has not learned to predict them. We propose coverage-aware reward calibration to use these outcomes as additional supervision. Before critic training, we fit a shared state potential to return-to-go labels on a nearest-neighbor graph, with local geometry controlling each label's influence. During training, we add the potential difference to the reward, amplify the original reward component, and gradually reduce both changes while retaining the base learner's update rules. Under the stated terminal convention, potential-based shaping preserves optimal action ordering at fixed coefficients. We relate potential error to behavior-advantage information and characterize when calibrated weights improve on uniform smoothing within locally constant-value regions. A two-step analytical example connects this improvement to the first Bellman update. Experiments cover sparse navigation, dexterous manipulation, and continuous control. Matched ablations show gains beyond reward scaling alone and support graph smoothing and geometry-based weighting, with additional controls across offline learners. These results support using precomputed outcomes to assist offline TD learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.