Value-guided Implicit Maximum Likelihood Estimation for Offline Reinforcement Learning
Abstract
Offline reinforcement learning aims to learn effective policies from pre-collected datasets without further environment interaction. However, offline datasets often exhibit multimodal behavior distributions, making accurate behavior modeling difficult and increasing the risk of selecting out-of-distribution actions during policy improvement. Recent diffusion- and flow-based policies improve multimodal modeling but often rely on iterative generation, distillation, or auxiliary optimization, leading to substantial computational overhead. To address these issues, we propose Value-guided Implicit Maximum Likelihood Estimation (VIMLE), an efficient offline RL framework that enables multimodal behavior modeling, behavioral consistency filtering, and value-guided action selection. Specifically, VIMLE uses multiple conditional IMLE generators for one-step parallel action generation, filters unreliable candidates through ensemble agreement, and selects high-value actions using an independently learned value function. This decoupled design enables policy improvement over behavior-consistent candidates while avoiding multi-step generation and auxiliary policy training. Extensive evaluations on D4RL demonstrate competitive performance across diverse tasks, while substantially reducing training time and inference latency compared with diffusion-based approaches. Code is provided in the supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.