acceptodds
Under review as a conference paper at ICLR 2027

Reasoning Canvas: Spatial-Temporal Visual Reasoning for Embodied Video Generation

Abstract

Video generation models are increasingly used as visual world models for embodied AI, but they are not explicitly trained for spatial-temporal reasoning required by multi-object, multi-step manipulation tasks. Such models may generate visually plausible videos while manipulating the wrong object, misordering subtasks, or placing objects at incorrect targets. To this end, we propose Reasoning Canvas, a dual-stream video diffusion framework that introduces an explicit auxiliary Thought stream for spatial-temporal structure. The Thought stream renders object-following bounding boxes and step-wise text prompts on a black canvas, providing a lightweight representation of what should happen before high-fidelity pixels are fully synthesized. A Timestep-Bias Reasoning Routing (TBRR) schedule resolves this simpler Thought stream earlier than the RGB stream, enabling subsequent RGB denoising to condition on a more stable plan. We also construct HierCanvas Dataset, a RoboCasa-derived hierarchical manipulation dataset with task→subtask→step programs, bounding boxes, and rendered step prompts. Experiments show that Reasoning Canvas yields consistent gains in visual quality, instruction alignment, and tracking, with stronger robustness to ambiguous instructions and OOD task compositions, validating ahead-of-time visual reasoning as an effective interface for video generation based world models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.