acceptodds
Under review as a conference paper at ICLR 2027

Q Learning with Normalizing Flows

Abstract

Q-learning algorithms are the dominant approach for reinforcement learning in discrete action spaces, but adapting them for continuous action spaces requires an intractable continuous max to find optimal actions. Maximum entropy reinforcement learning replaces the max with a softmax that allocates action density proportionally to the exponentiated Q-function, but without an efficient to maximize the Q-function this alone does not suffice for continuous Q-learning. This paper introduces Q Normalizing Flows (QNF), a class of methods that directly represent the softmax density via normalizing flows, allowing effective semi-gradient Q-learning in continuous action spaces. This makes a rich body of representations immediately available, including those most commonly used in off-policy actor critic methods, without coupling heavily to specialized architectures for action selection. Q Normalizing Flows perform similarly to widely used actor-critic algorithms for continuous RL while training with the pure semi-gradient action evaluation loss.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.