Bellman Calibration for V-Learning in Offline Reinforcement Learning
Abstract
Calibration is a standard tool for assessing and correcting the reliability of predictions in supervised learning, but is less developed for long-horizon value prediction in reinforcement learning, where the prediction target is not directly observed. We extend calibration to the Bellman fixed-point setting by introducing Bellman calibration, which requires the Bellman equation to hold conditionally on the predicted value. Equivalently, among states assigned the same predicted value, the average one-step Bellman target should equal that prediction. Because this target itself depends on the value predictor, calibration becomes a fixed-point problem. We define a calibration error that quantifies systematic Bellman residuals conditional on the predicted value and develop doubly robust estimators from off-policy data. We establish a calibration–refinement bound that relates calibration error to value-prediction error, separating the component that post-hoc calibration can correct from information lost by the original predictor. We then develop Fitted Bellman Calibration, a model-agnostic post-hoc procedure that iteratively recalibrates a learned value predictor using one-dimensional histogram or isotonic regression, without refitting the underlying high-dimensional value function. We establish finite-sample guarantees showing one-dimensional nonparametric calibration rates for histogram and isotonic calibration, together with value-error guarantees for fixed-partition calibration, without imposing realizability or Bellman-completeness assumptions on the calibrated predictor class.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.