acceptodds
Under review as a conference paper at ICLR 2027

The Critic Is Only a Proxy: Regret Optimization for Reinforcement Learning

Abstract

When direct evaluation is costly, learning systems can turn limited observations into a reusable source of feedback. Actor-critic reinforcement learning follows this approach, but its critic is only an informative, imperfect proxy for environmental return. The actor therefore faces an intrinsic robustness problem even without environment shift. We propose to protect against the worst plausible regret: the gap between the current policy and the best action under the same plausible critic. To develop a practical algorithm, we further replace hard hindsight search by a soft comparison and develop Distributionally Robust Relaxed Regret Optimization (DR3O), which can be efficiently estimated from counterfactual action samples. This principle can serve as a general framework when setting up the objective against which the actor optimizes. As an example, we use this approach to indicate "which to give to the actor" in the well-known twin critics. As a result, across MuJoCo, MetaWorld, and the DeepMind Control Suite, changing this actor-facing signal improves matched training backbones without any additional efforts. A data-dependent ellipsoidal bonus provides further gains, supporting regret optimization as a practical way to learn from imperfect feedback.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.