DELTA: Dual-stream Encoding with Latent Temporal Anticipation for Frozen Visual World Models
Abstract
Frozen vision backbones such as V-JEPA 2.1 provide strong visual features for manipulation, but not an action-conditioned model of dynamics. Existing world models introduce that dynamics under different modeling assumptions. Predictors defined on frozen features take a short temporal window and regress the next embedding, so they maintain no state beyond that window and do not cast the prediction as an update of the observed frame. Predictors that retain a longer history instead advance an image code or a learned latent, and therefore depart from the frozen representation. We propose DELTA, a reconstruction-free world model on frozen V-JEPA 2.1 (ViT-B/384) features. A dual-stream latent carries state across time and is updated by an action-conditioned causal transition. A residual head writes the resulting short-horizon change onto the observed tokens. Actions enter only as conditioning. On manipulation demonstrations, under a shared feature-cache protocol, DELTA lowers multi-step open-loop, replan, and CEM feature-prediction error relative to these baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.