acceptodds
Under review as a conference paper at ICLR 2027

Regulating Actor-Critic Coupling with Uncertainty in Offline Reinforcement Learning

Abstract

Offline reinforcement learning (RL) is particularly vulnerable to extrapolation errors induced by out-of-distribution (OOD) actions. While existing methods typically address this problem through policy constraints or conservative value estimation, we argue that the key difficulty in offline actor-critic learning lies in how estimation errors are amplified through the coupling between policy improvement and boostrapped value evaluation. In this work, we introduce Uncertainty-Regulated Actor-Critic (URAC), which uses uncertainty as a shared reliability signal to regulate both stages of this coupling. URAC learns a data-conditioned value anchor together with an uncertainty scale. On the critic side, it suppresses uncertain optimistic bootstrap values before they propagate through Bellman updates; on the actor side, it adaptively controls how much value-guided residual deviation from a behavior anchor is accepted. We further show that these two mechanisms provide complementary contractions of actor-side error admission and critic-side error transmission, yielding a multiplicative reduction in closed-loop error propagation. Across Gym-MuJoCo and NeoRL2 benchmarks, URAC achieves the highest aggregate performance among the evaluated methods, while remaining robust when offline data are limited or corrupted by noise. These results suggest that explicitly regulating actor-critic coupling provides an effective way to stabilize policy improvement under uncertain offline supervision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.