acceptodds
Under review as a conference paper at ICLR 2027

SNARL: Self-Normalizing Anchored Maximum-Entropy Reinforcement Learning

Abstract

Expressive generative policies can capture complex action distributions, but maximum-entropy (MaxEnt) reinforcement learning requires evaluating the policy log-density , which is generally intractable for implicit generators. We introduce Self-Normalizing Anchored MaxEnt Reinforcement Learning (SNARL), an off-policy framework that learns a separate state-conditioned energy model to estimate the policy's negative log-density from samples. SNARL augments spatial and temporal score matching with an anchor derived from the Gaussian limit of variance-preserving corruption, approximately normalizing the energy model as the policy evolves. The resulting model estimates the policy's negative log-density in a single forward pass. By separating policy log-density estimation from action generation, SNARL enables maximum-entropy learning with a one-step pushforward actor that generates each action in a single network evaluation. Experiments show that SNARL is competitive with strong diffusion and flow baselines across high-dimensional continuous-control benchmarks, excels on MuJoCo Playground humanoid locomotion tasks, and provides approximately faster policy inference than DIME.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.