acceptodds
Under review as a conference paper at ICLR 2027

Selective Factorized Risk Regularization for Offline Reinforcement Learning

Abstract

Offline reinforcement learning must balance policy improvement against unreliable extrapolation beyond the available data. Local action support and marginal successor-state support provide useful evidence, but do not establish whether a specific state–action–successor relation is supported or whether subsequent policy decisions will remain supported. Moreover, insufficient data support does not necessarily imply low return, motivating the use of value information when determining how support evidence should constrain learning. We propose Selective Factorized Path Risk (SFPR), a framework that regularizes policy improvement through factorized, value-guided future support costs. SFPR separately models action support conditioned on the current state and transition support conditioned on the state–action pair. Two branch-specific cost estimators recursively connect these local assessments to subsequent decisions. Data-support and behavior-relative value assessments regulate batch-level cost strengths, while support-dependent continuation reduces recursive propagation at weakly supported transitions. SFPR integrates these future costs with behavior-distribution fitting and selective value regularization to constrain both policy updates and the value estimates guiding them. SFPR achieves the highest average performance among the compared methods on both the D4RL locomotion and Maze2D navigation benchmarks. These results demonstrate SFPR's overall performance advantages across continuous-control and navigation tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.