acceptodds
Under review as a conference paper at ICLR 2027

Typicality Constrained Latent Reinforcement Learning for Generative Policies

Abstract

Diffusion steering via reinforcement learning (DSRL) improves a frozen diffusion or flow-matching policy by learning a latent policy over its input noise. The frozen policy, however, acts well only on noise typical of its Gaussian prior, with a tolerance that shrinks as the dimension grows. Existing steering methods bound the noise in raw units re-tuned per task and need a costly critic distilled through the frozen policy. We propose Typical DSRL, whose latent policy is the prior shifted by a learned mean, so its KL divergence from the prior is exact; bounding it in widths of the prior's typical shell lets one budget serve every dimension. Because its critic sees only typical noise, it needs no distillation. On eleven robomimic, Adroit and Gym tasks, with one budget, Typical DSRL is Pareto-optimal in performance against training time; baselines that come close train two to twenty-four times longer. Typicality is needed only in training: deploying the latent policy's mean, which lies far from the typical shell, outperforms sampling on every task.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.