T-LESSA: TEMPORAL-LOGIC EVENT SPECIFICATION STEERING OF VIDEO DIFFUSION ATTENTION
Abstract
Recent advancements in generative video promise to transform challenging scenarios into training data: rare driving situations for autonomous vehicles, synthetic scenes for robot simulation, and realistic deepfakes for stress testing detection models. However, in practice, we still prompt a model and pray to get faithful videos. While a large body of research exists on generating good video, there is a lack of methods to enforce temporal and logical requirements (e.g., ordering, negation, co-occurrence) during video generation. Prior work has focused on fine-tuning with preference feedback or verifiable rewards. However, these methods require additional training and reduce a temporal-logic requirement to a scalar reward. This limitation has motivated a recent line of work that steers video generation at inference time without any training. Although these approaches control when each required event appears, they do not ensure that the generated videos align with their temporal and logical requirements. Therefore, we introduce *T-LESSA*, a training-free method that steers video generation toward satisfying a temporal-logic requirement. We show that *T-LESSA* improves temporal completion by **14.3%** and, across four backbones, temporal-logic satisfaction by **21.8%** over the unsteered baseline, at 1.7% to 13.5% additional generation time. Furthermore, *T-LESSA* improves compositional alignment in 14 of 16 backbone-category pairs on T2V-CompBench, and largely preserves video quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.