acceptodds
Under review as a conference paper at ICLR 2027

Calibrated Pessimism in Offline Reinforcement Learning: Auditing Support-Conditioned Reliability and Control Transfer

Abstract

Offline reinforcement learning often relies on pessimism to reduce the influence of value estimates in weakly supported regions, yet the uncertainty signals used to construct pessimistic penalties are rarely tested as calibrated error surrogates. We study this reliability question directly in a focused five-seed HalfCheetah Medium case study. Using leakage-controlled train/calibration/validation splits, we calibrate global, uncertainty-ratio, and local-support-aware bounds for an observable held-out Bellman-residual proxy, and separately evaluate marginal coverage, support-conditioned reliability, and sharpness. Calibrated constructions attain near-nominal marginal coverage, but this aggregate view hides substantial redistribution of reliability across support regimes: at 90% nominal coverage, the global bound remains relatively narrow while under-covering low-support transitions, whereas adaptive scaling raises low-support coverage at the cost of substantially wider bounds and lower high-support coverage. We then reuse the frozen calibration objects in an evaluation-time candidate-reranking study with matched proposals, no retraining, and no return-based tuning. The reranking interface itself substantially underperforms the deterministic actor, and within this restricted interface we find no evidence that calibrated penalties improve return over unpenalized Q reranking. Our results therefore show, within the scope of this case study, that uncertainty magnitude, marginal statistical calibration, and downstream control utility can diverge substantially. The standard split-conformal guarantee applied here concerns only marginal coverage of the observable residual proxy under exchangeability assumptions; it is not a guarantee for true Bellman error, policy value, or return.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.