Paired Relative-Value Regularization for Offline Reinforcement Learning
Abstract
Offline reinforcement learning (RL) suffers from unreliable value estimates for actions that are insufficiently covered by the offline dataset. Conservative value learning mitigates this issue by suppressing the values of such actions, but overly strong regularization can also limit useful policy improvement. In this work, we propose Paired Relative-Value Regularization (PaRE), which approaches conservative value regularization from a policy-local perspective. PaRE constructs candidate actions around the current policy and regularizes their values relative to the data action observed at the same state, allowing the regularization strength to adapt to the local relative-value structure. We theoretically analyze the relationship between PaRE and policy-local value overestimation, and characterize how its gradient determines the overall regularization strength and its allocation across candidates. Experiments on the D4RL MuJoCo locomotion and Adroit dexterous manipulation benchmarks show that PaRE achieves competitive performance across datasets with different data qualities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.