Flow Q-trace: Multi-step Critic Learning with Density-Ratio-Based OOD Action Handling
Abstract
In this paper, we consider multi-step critic learning for offline reinforcement learning (RL), moving beyond the widely-used one-step temporal difference methods. Although multi-step learning can offer more accurate value estimation, it has been underutilized in offline RL settings due to two challenges: the difficulty of correcting multi-step off-policyness via importance sampling, and value overestimation caused by bootstrapping from out-of-distribution (OOD) actions. We address these challenges by introducing a conditional density ratio estimator and develop Q-trace, an offline multi-step critic learning algorithm with several desirable properties. In our framework, the learned density ratio estimator is employed not only for off-policyness correction but also for effective OOD action handling through critic penalization and rejection sampling, exploiting its capacity to identify OOD actions in a pointwise manner. By integrating Q-trace with expressive flow-based policies and our density-ratio-based OOD handling, we propose Flow Q-trace (FQT), which demonstrates strong performance across diverse offline RL tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.