WorldCeption: Unifying Perception and Generation for World Modeling
Abstract
Visual world modeling requires both inferring structure from observations and generating visual predictions under structural conditions. Although generative priors support visual perception, accurate perceptual readout alone does not establish that its representations can guide synthesis. A central challenge is to make dense perceptual fields and sparse states serve both purposes within a unified model. We introduce WorldCeption, a compact framework that unifies perception and conditional generation through a shared dense–sparse interface. Given RGB video, WorldCeption predicts dense perceptual representations and continuous or discrete sparse states; conversely, it synthesizes RGB video conditioned on dense representations or sparse states. Built upon a pretrained video transformer, the model combines shared attention with role-dependent feed-forward computation and task-dependent information access, accommodating perceptual readout and conditional synthesis within the same network. These complementary mappings can be composed, enabling estimated states to guide generation and perceptual representations to be translated through an RGB intermediate. Experiments show that estimated camera states retain reconstruction quality close to annotated conditions, while changing supplied states redirects recovered camera motion. Composed perceptual and generative paths further enable cross-representation completion without direct pairwise supervision. WorldCeption offers a compact architectural extension toward perception-centric world modeling through reusable visual representations. Code and models will be available upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.