Learning Maximally Computable Visual Representations for Controllable Generation
Abstract
3D reconstruction provides explicit geometry and accurate transformation of observed content, but requires dense inputs and often degrades under large viewpoint changes. In contrast, generative world models can synthesize realistic future observations, yet their data-driven dynamics make precise geometric control and long-horizon consistency difficult to guarantee. We introduce Maximally Computable Visual Representation (MCVR), which combines explicit reconstruction with generative modeling. Rather than predicting the entire future visual state, MCVR learns a persistent representation whose physically supported content is deterministically propagated under known camera and object transformations, while learned generation is reserved for unrecoverable content. This enables accurate object manipulation and geometry-aware generation while reducing unnecessary prediction. Experiments on autonomous-driving scenes show that MCVR achieves a strong trade-off between visual sufficiency and physical computability, improving controllability and geometric consistency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.