World in World: Explore the World with World Models
Abstract
Recent advances in video world models enable interactive generation beyond fixed-length synthesis, yet extending pretrained models to new forms of control typically requires control-specific architectures or additional training. We introduce World in World (WiW), a training-free visual-evidence interface that flexibly extends the controllability of frozen causal video world models. WiW leverages native self-attention, which already consumes clean visual states, as a shared interface for heterogeneous external controls without modifying the pretrained backbone. It converts source observations, target-view projections, rendered geometry, and generated history into clean visual evidence annotated with camera, temporal, and spatial-validity information. These sources provide appearance references, spatial guidance, completion cues for newly visible regions, and long-range context for revisiting generated states. We introduce correspondence-guided attention routing (CGAR) to direct queries toward geometrically corresponding visual tokens, and evidence-wise attention CFG (EWA) to regulate the influence of different evidence sources within native self-attention without additional network function evaluations for guidance. We instantiate WiW on a frozen LingBot-World 2.0 backbone and evaluate camera-controlled video rerendering on DAVIS and OpenVid-1M across diverse viewpoint changes. Beyond rerendering, WiW enables exploration within a given video's dynamic world and supports bullet-time generation, video stabilization, video editing, and cross-generation K/V sharing, demonstrating a unified and versatile framework for training-free world-model control.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.