acceptodds
Under review as a conference paper at ICLR 2027

ChronoGuide: Training-Free Temporal Control for Video Generation via Cross-State Attention

Abstract

Despite major gains in visual quality, text-to-video models remain unreliable at precise temporal control. Objects may appear early or late, persist too long, or fail to return on time, and actions may peak off cue. We identify a structural source of these failures in spatiotemporal self-attention. The same cross-frame communication that supports coherence also transfers features between frames requiring different states, such as presence and absence, pulling them away from the requested schedule. ChronoGuide addresses this conflict with a training-free cross-state attention mechanism. Its core operator preserves live attention weights and same-state contributions, replacing only values crossing state boundaries with matched features computed under the receiving frame's requested state. This targets cross-state interference while retaining same-state communication, without training or iterative latent optimization. We further introduce a mixed-lifecycle object evaluation covering entry, exit, bounded presence, and return, and a complementary evaluation using an open-vocabulary detector. Across all benchmark panels evaluated against TempoControl, ChronoGuide improves every reported metric. On the standard 80-prompt one-object Wan benchmark, it reaches 90.06% temporal accuracy, 83.88% presence accuracy, and 96.25% absence accuracy versus TempoControl's 83.56%, 79.75%, and 87.38%. It also uses substantially less peak VRAM than TempoControl, enabling deployment on lower-memory GPUs and batched inference. ChronoGuide transfers across continuous and discrete video diffusion models, including CogVideoX and URSA.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.