acceptodds
Under review as a conference paper at ICLR 2027

MALA-Guided Mirror Descent

Abstract

Policy mirror descent (PMD) defines policy improvement as a simple distribution-space update: exponentially tilt the current policy toward high-value actions while remaining close to the original policy. For expressive generative policies, however, realizing this prescribed target distribution can itself be difficult. We introduce MALA-Guided Mirror Descent (MGMD), which treats policy improvement as an inference problem rather than solely as an actor-training objective. MGMD represents the diffusion policy through an energy function and uses annealed Metropolis-adjusted Langevin dynamics to traverse a sequence of value-tilted noisy distributions whose zero-noise limit is the PMD target. The resulting improved actions are used for environment interaction and Bellman targets, and are continuously distilled back into the diffusion policy with the standard denoising objective. On a controlled multimodal bandit where the exact PMD target is known, MGMD tracks the target more accurately than advantage-weighted regression as the target moves farther from the base policy. Across six MuJoCo continuous-control tasks, MGMD demonstrates consistently strong performance. These results suggest that inference-time computation can provide a persistent mechanism for policy improvement when it is directed toward a prescribed policy-improvement distribution and subsequently amortized into the policy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.