acceptodds
Under review as a conference paper at ICLR 2027

Efficient Implicit Boltzmann Policy Optimization via Mixture Likelihood Maximization

Abstract

Online reinforcement learning often requires representing complex, potentially multimodal action distributions induced by the value function. Recent diffusion and flow-based policies have substantially improved policy expressivity by enabling flexible approximations to complex action distributions. However, it remains challenging to achieve both high sampling fidelity and low inference latency with expressive generative policies, as they typically generate actions iteratively. While recent one-step approaches reduce inference cost, training these expressive policies has remained computationally expensive. We introduce **Efficient Implicit Boltzmann Policy Optimization via Mixture Likelihood Maximization (iBOLT)**, which efficiently trains a one-step policy to approximate the value-induced Boltzmann distribution. iBOLT represents the policy as a continuous mixture of truncated Gaussians, with component parameters generated from the state and a latent variable by a single network. Then, we use a tractable finite-mixture approximation to reweight sampled actions toward the Boltzmann target and train the policy by maximizing the weighted mixture likelihood. During deployment, the learned policy generates actions with a single network evaluation, without iterative refinement. Experiments on MuJoCo benchmarks and multimodal tasks demonstrate that iBOLT achieves performance competitive with state-of-the-art generative policy methods, while substantially reducing both training cost and inference latency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.