DEFlow: Deep Equilibrium Flow Maximum-Entropy Reinforcement Learning
Abstract
Expressive flow policies can represent multiple high-value actions, but sampling them with few integration steps while retaining tractable action densities remains challenging. We introduce DEFlow, a deep equilibrium flow policy that replaces explicit Euler sampling with implicit Euler integration and solves all intermediate actions jointly as an equilibrium system. We combine implicit differentiation with a discrete change-of-variables formula for the final-action density, enabling off-policy training within soft actor-critic without backpropagating through solver iterations. DEFlow achieves an average relative improvement of approximately 10% in mean successful coverage over the strongest baseline across four Multi-Goal docking variants. It also achieves the highest mean time-averaged return on four of five manipulation tasks, while remaining broadly competitive across five MuJoCo tasks. These results support implicit flow policies for tasks requiring both diversity and precision, at the cost of additional nonlinear and linear solves.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.