SUNDER: A General Framework for Decoupled State–Policy Optimization in Offline Reinforcement Learning
Abstract
Distribution Correction Estimation (DICE) methods directly optimize policy-induced state-action occupancies, enabling distributional objectives beyond standard behavior-regularized reinforcement learning. However, joint occupancy regularization couples state visitation and action selection through a single regularization strength, even though these two forms of distribution shift may require different control. Naively separating them breaks the explicit affine structure underlying the standard DICE strong-duality argument. We introduce SUNDER (State–policy UNcoupling for Distribution Estimation via Ratio factorization), a general framework that separately controls state visitation and policy deviation. Although the resulting state–policy formulation is non-affine, we show that it admits a hidden convex structure that preserves strong duality and enables tractable elimination of the explicit optimization variables. We instantiate SUNDER for reward maximization, state-entropy maximization, and learning from observations. Under mismatches between state coverage and action quality, Opt-SUNDER outperforms OptiDICE across six maze domains, while achieving the best average performance among competitive recent baselines on 27 Minari tasks. Ent-SUNDER consistently outperforms SEMDICE across seven maze and four DMC domains with skewed dataset state distributions, and LfO-SUNDER achieves competitive performance on MuJoCo and Adroit tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.