acceptodds
Under review as a conference paper at ICLR 2027

WOVEN: Weaving Visual World Modeling into Multimodal LLMs

Abstract

Multimodal large language models (MLLMs) exhibit scattered failures across spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in action-conditioned visual transition reasoning over triplets, and investigate its internal organization to derive a systematic training recipe. To enable controlled study, we introduce WOVEN, a benchmark and training source that factorizes transition supervision along independently controllable semantic dimensions using diverse, realistic rollouts from video-pretrained generative models. WOVEN reveals a substantial gap to human performance, motivating a direct study of trainability and transfer. Across controlled training subsets of only approximately 2,000 examples, post-training improves 22 of 26 external benchmarks, with gains up to 27.3 percentage points. Positive transfer persists across scales, and individual downstream tasks also benefit from distinct supervision sources; WOVEN examples can further replace part of task-specific training data while retaining comparable accuracy, establishing transition reasoning as a shared, trainable primitive. More importantly, controlled comparisons yield a training recipe organized by the query operator over : curricula sharing an operator exhibit consistently similar downstream gain profiles, whereas shared actions, scenes, or domains do not explain this grouping. Additional findings connect broader state interventions to robustness and delineate complementary supervision needs and transfer boundaries. Our work establishes visual transition reasoning as a reusable foundation for improving diverse downstream capabilities, paving the way for systematic visual world-model training guided by the structure of transfer.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.