acceptodds
Under review as a conference paper at ICLR 2027

Why Do Text-to-Video Models Struggle with Spatial Relations?

Abstract

Text-to-video (T2V) models have achieved impressive generation quality, yet they remain unreliable in establishing spatial relations among multiple objects. In this work, we investigate the sources of spatial-relation failures in transformer-based T2V models. Our analysis identifies two key limitations. First, the frozen text encoder weakly represents the implicit complementary spatial role of reference objects. Second, joint conditioning produces ambiguous object grounding during early denoising. Based on these observations, we propose Spatial Patching and Anchored Relational Conditioning (SPARC), a training-free approach for improving spatial controllability directly from text prompts. Our method first applies Complementary Spatial Token Patching to strengthen the spatial representation of reference objects. It then uses Anchor-to-Graph Denoising to progressively establish relations around a shared spatial anchor. Both components operate at inference time with frozen model parameters and do not require additional spatial inputs. Experiments across two T2V backbones and multiple spatial benchmarks demonstrate consistent improvements in spatial controllability. As a training-free method, SPARC achieves a 12.5% relative improvement in the T2V-CompBench spatial score over the second-best method.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.