acceptodds
Under review as a conference paper at ICLR 2027

Laplacian Flows for Policy Learning

Abstract

Modern AI systems improve task performance through local policy updates, yet gains on the current training objective can come at the expense of previously acquired behaviours, a risk that conservative single-step updates alone do not eliminate. To guide adaptation with a reference to accumulated experience, we draw on geometric measure theory and model historical policies as samples from an experience-induced, locally low-dimensional manifold in Wasserstein space. We introduce Policy Laplacian Trace (PLT), an evolving trajectory map that uses geometric relations among relevant historical policies to regularise subsequent updates. Unlike Wasserstein constraints imposed only between successive iterates, PLT relates each candidate update to multiple historical policies. A history-induced graph Dirichlet energy translates these geometric relations into a Laplacian-type correction to the task-driven update, bringing trajectory-level geometric context into local optimisation. Learning shapes the map, while its evolving graph geometry guides subsequent policy updates. We instantiate PLT in reinforcement learning and language-model training, using a scalable semantic-transport proxy for the latter. Experiments with PPO and MAPPO on Atari, MuJoCo, and StarCraft II show improvements in final performance and sample efficiency. Controlled language-model stress tests show gains in robustness to contradictory evidence, long-range factual recall, and few-shot generalisation. On the recall task, Qwen3-4B improves weighted token accuracy by 4.83%, with a 3.56% increase in median training-step time under the tested configuration. This feedback between policy evolution and its geometric map suggests a potential role for experience-induced geometry in recursive learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.