acceptodds
Under review as a conference paper at ICLR 2027

Learning to Optimize Actions with Flows

Abstract

Model-based planning requires repeatedly solving complex non-convex optimization problems over sequences of actions. These problems are predominantly tackled by sampling-based optimizers, whose hand-crafted update rules cannot exploit structure patterns across planning instances. From a complementary perspective, generative modeling has enabled expressive policies that directly generate action sequences, yet without optimizing them over a horizon through a predictive model. We bridge these two paradigms by introducing **FlowOpt Policy**, which casts planning as a learned transport problem over action sequences: rather than generating actions directly, a flow iteratively transports a population of candidate action sequences toward lower cost, conditioned on their rollout evaluations. The flow is parameterized by an attention mechanism in which every candidate attends over the evaluated population, so that its update depends jointly on where the other candidates lie and how they scored. We further reveal connections to classical methods such as MPPI and CEM, whose updates emerge as particular fixed attention patterns. FlowOpt Policy is trained in a self-supervised manner via flow matching on improvement steps computed by a sampling-based optimizer on the populations the planner encounters, without requiring any demonstration data. Experiments on classical robotics and world model planning tasks show that FlowOpt Policy consistently outperforms sampling-based planners at equal rollout budget, and is more robust than state-conditioned generative policies, generalizing better to unseen objectives because it learns how to optimize rather than what to output.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.