acceptodds
Under review as a conference paper at ICLR 2027

Action Jets: Charting Visual Action Manifolds with Generative Oracles

Abstract

World models are a powerful means of reasoning and planning for a variety of applications including robotics. However, many of these world models such as video generation models are high dimensional and computationally expensive which limits the number of action candidates that can be evaluated by a planner. We show that this cost is avoidable for planners over short horizons which only local information with respect to the current state. For a fixed state, the frames reachable by varying a low-dimensional action, a camera movement or an end-effector command, lie on a low-dimensional manifold. We estimate the local shape of that manifold directly via a network trained to estimate how each image region moves per unit of each action axis supervised by data collected from samples by the generator. A correction head fit to trajectory data estimates the shape of this manifold region up to a third-order Taylor polynomial. We call this network estimating the Taylor polynomial describing the manifold shape an action jet and compute its radius of validity in closed form. By traversing this approximation of the local manifold shape, we only have to perform inference once per state for any number of candidate actions. We achieve over 100x speedup of same length trajectory rollouts compared to the source generators, and this gap widens with the number of candidates a planner evaluates. We measure against state-of-the-art action-conditioned generators: IRASim and WorldGym on BridgeData V2, a real 7-DoF manipulation benchmark, and DepthSplat, LagerNVS and FrameCrafter on DL3DV scenes. Because only the first-order term comes from the generator and the rest from the data, the action jet also predicts held-out frames more accurately than IRASim, the generator that supplied its tangent, at short horizons and at roughly x less compute.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.