Beckmann World Models: Direct Terminal-Map Learning for One-Step Action-Conditioned Prediction
Abstract
Action-conditioned world models must predict the consequences of proposed controls quickly enough to support repeated queries. We introduce Beckmann World Models (BWM), which learn a direct map from noise, visual history, and actions to future RGB frames without a pretrained generative teacher. Beckmann transport supplies a conservation-based training rule for this terminal map using sampled noise–future pairs, without a population of generated negatives. However, learning the map on interpolated inputs does not ensure accurate prediction from the pure noise used at deployment. We rectify this source-endpoint gap by proposing endpoint-aligned training: direct supervision of the source query, motion and temporal losses, and spatial feature supervision, supported by a generated-history regularizer. These additions train the same RGB map and preserve one network call per future chunk. On PushT and Robomimic Can, BWM reduces 64-frame MSE by – relative to same-backbone regression at comparable inference cost, with measured latency of – ms per native chunk. It also improves prediction over locally trained DriftWorld and policy-performance estimation, while the planning comparison shows no detected improvement. Cumulative training costs are disclosed because these are comparisons of complete recipes. The results identify a practical route from autonomous transport to accurate one-call world prediction. Anonymous project page: https://beckmann-world.pages.dev/https://beckmann-world.pages.dev/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.