Typicality Constrained Latent Reinforcement Learning for Generative Policies
Abstract
Diffusion steering via reinforcement learning (DSRL) improves a frozen diffusion or flow-matching policy by learning a latent policy over its input noise. The frozen policy, however, acts well only on noise typical of its Gaussian prior, with a tolerance that shrinks as the dimension grows. Existing steering methods bound the noise in raw units re-tuned per task and need a costly critic distilled through the frozen policy. We propose Typical DSRL, whose latent policy is the prior shifted by a learned mean, so its KL divergence from the prior is exact; bounding it in widths of the prior's typical shell lets one budget serve every dimension. Because its critic sees only typical noise, it needs no distillation. On eleven robomimic, Adroit and Gym tasks, with one budget, Typical DSRL is Pareto-optimal in performance against training time; baselines that come close train two to twenty-four times longer. Typicality is needed only in training: deploying the latent policy's mean, which lies far from the typical shell, outperforms sampling on every task.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.