Any-to-Any Video Modeling: Modeling the World by Traversing Time and Modalities
Abstract
Video modeling captures how the world is structured in space and how it evolves over time. Most approaches model this structure in a single representation space: RGB or a semantic representation (e.g., DINOv2, V-JEPA). However, different spaces capture different aspects of the underlying world and serve complementary roles, e.g., DINOv2 and V-JEPA excel at high-level semantic understanding, while RGB preserves fine-grained detail. In this work, we move beyond modeling the world in a single representation space and study video modeling jointly across time and modalities, each serving as a representation space (e.g., RGB, depth, optical flow, feature maps, and bounding boxes etc.). We develop an any-to-any multimodal video model that can predict any modality at any point in time, given any combination of input modalities at any time steps. Flexibly traversing this time-modality grid gives rise to higher-level inference capabilities that are not supported by specialist video models. First, in chained generation, we decompose a task (e.g., caption-to-video) into a chain of potentially simpler sub-tasks of predicting intermediate modalities, and find that this yields higher-quality generations than direct generation. Second, we study future prediction in different modality spaces and find that no single space is optimal: the best prediction space varies across samples, and selecting it per sample yields substantially more accurate forecasts than committing to any single space. Finally, we show that the flexible modality-time conditioning enables counterfactual generation, where conditioning modalities act as actions that steer generation, from coarse (e.g., captions) to fine-grained (e.g., depth or bounding boxes) control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.