Reasoning Canvas: Spatial-Temporal Visual Reasoning for Embodied Video Generation
Abstract
Video generation models are increasingly used as visual world models for embodied AI, but they are not explicitly trained for spatial-temporal reasoning required by multi-object, multi-step manipulation tasks. Such models may generate visually plausible videos while manipulating the wrong object, misordering subtasks, or placing objects at incorrect targets. To this end, we propose Reasoning Canvas, a dual-stream video diffusion framework that introduces an explicit auxiliary Thought stream for spatial-temporal structure. The Thought stream renders object-following bounding boxes and step-wise text prompts on a black canvas, providing a lightweight representation of what should happen before high-fidelity pixels are fully synthesized. A Timestep-Bias Reasoning Routing (TBRR) schedule resolves this simpler Thought stream earlier than the RGB stream, enabling subsequent RGB denoising to condition on a more stable plan. We also construct HierCanvas Dataset, a RoboCasa-derived hierarchical manipulation dataset with task→subtask→step programs, bounding boxes, and rendered step prompts. Experiments show that Reasoning Canvas yields consistent gains in visual quality, instruction alignment, and tracking, with stronger robustness to ambiguous instructions and OOD task compositions, validating ahead-of-time visual reasoning as an effective interface for video generation based world models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.