Efficient Implicit Boltzmann Policy Optimization via Mixture Likelihood Maximization
Abstract
Online reinforcement learning often requires representing complex, potentially multimodal action distributions induced by the value function. Recent diffusion and flow-based policies have substantially improved policy expressivity by enabling flexible approximations to complex action distributions. However, it remains challenging to achieve both high sampling fidelity and low inference latency with expressive generative policies, as they typically generate actions iteratively. While recent one-step approaches reduce inference cost, training these expressive policies has remained computationally expensive. We introduce **Efficient Implicit Boltzmann Policy Optimization via Mixture Likelihood Maximization (iBOLT)**, which efficiently trains a one-step policy to approximate the value-induced Boltzmann distribution. iBOLT represents the policy as a continuous mixture of truncated Gaussians, with component parameters generated from the state and a latent variable by a single network. Then, we use a tractable finite-mixture approximation to reweight sampled actions toward the Boltzmann target and train the policy by maximizing the weighted mixture likelihood. During deployment, the learned policy generates actions with a single network evaluation, without iterative refinement. Experiments on MuJoCo benchmarks and multimodal tasks demonstrate that iBOLT achieves performance competitive with state-of-the-art generative policy methods, while substantially reducing both training cost and inference latency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.