Video Models Reason Early: Exploiting Plan Commitment for Maze Solving
Abstract
Video diffusion models exhibit reasoning capabilities beyond content creation like solving mazes and puzzles, yet little is understood about *how* they reason during generation. We take a first step towards understanding this and study the internal planning dynamics of video models using 2D maze-solving as a controlled testbed. Our investigations reveal two findings. Our first finding is **early plan commitment**: video diffusion models commit to a high-level motion plan within the first few denoising steps, after which further denoising alters visual details but not the underlying trajectory. Our second finding is that **path length**, not obstacle density, is the dominant predictor of maze difficulty, with a sharp failure threshold at 12 steps. Motivated by these findings, we introduce **Ch**aining with **Ea**rly **P**lanning, or ChEaP, an inference-time strategy that only spends compute on seeds with promising early plans and chains them together to tackle complex mazes. ChEaP improves accuracy from 7% to 67% on long-horizon mazes and by overall on hard tasks across Frozen Lake and VR-Bench using Wan2.2-14B and HunyuanVideo-1.5.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.