Video Reasoning Generalizes to Perception, Embodiment, and Real-World Simulation
Abstract
Progress in video generation has been driven largely by visual fidelity. Currently, these models are being deployed as world models and embodied planners, where success depends on reasoning-like competences rather than on pixels. Whether such competences can be acquired by training, and whether they generalize beyond the domain they were trained on, remains unanswered. We introduce XVReason and XVReason-Bench, a video corpus that pairs training data with matched evaluation across four domains: abstract visual reasoning, low-level perception, embodied manipulation, and world simulation. On this foundation, we run a cross-domain study on two video backbones. Generalization turns out to be asymmetric. Abstract reasoning is the only fine-tuning domain that improves all four task domains on both backbones, with gains of +42.3/+20.6/+10.7% on the three application domains of Wan2.2. This generalization ability also works in different visual domains. Follow-up experiments show that generalization varies with the adaptation regime, the data scale, the denoising stage, and the task composition, and trace the gain to three competences: instruction following, goal completion, and scene stability. An expert-loading ablation recovers most of the reasoning gain through the high-noise expert. These findings point to a recipe, ReasonGraft, which blends a reasoning LoRA with an in-domain LoRA at a 3:7 ratio in weight space. It beats in-domain fine-tuning on all three application domains. Video Reasoning is an under-used general-purpose resource for training video generation models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.