Fourier-World: Temporal Fourier States for Visual World Models
Abstract
World models aim to learn the dynamics of the world by predicting the future from visual observations. Recent approaches have shifted from directly generating RGB frames to predicting future features of vision foundation models (VFMs), which can support diverse downstream tasks. However, small feature errors accumulate during autoregressive prediction, degrading long-term rollout performance, while how existing predictors internally represent and use temporal changes remains unclear. We propose Fourier-World, which mitigates error accumulation by explicitly structuring temporal context as a frequency-domain state. Fourier-World applies a Fourier transform to the four most recent context features. A query initialized from the latest feature reads the resulting Fourier coefficients to predict feature-space deltas, which are then accumulated to generate future dense VFM features. Pretrained on the large-scale Kinetics-700 dataset, Fourier-World demonstrates strong transfer performance in segmentation, depth estimation, action anticipation, and causal future-point forecasting across five benchmarks: Cityscapes, VSPW, KITTI, EK100, and TAP-Vid-RGB-Stacking. It also maintains stable predictions over long autoregressive rollouts, achieving high computational efficiency and competitive performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.