acceptodds
Under review as a conference paper at ICLR 2027

On Time: Learning to Follow Event Timelines in Long Video Generation

Abstract

On-time long video generation requires actions to follow prescribed start and end times, preserve intended overlaps, and remain associated with the correct participants throughout an extended rollout. We introduce OnTime, a framework for this task that combines structured event timelines with a few-step autoregressive generator trained on short sequences. Its core mechanism, Rolling Timeline Conditioning (RTC), preserves global action intervals while assigning persistent scene and identity descriptions bounded local support, so their encoded durations do not grow with the full rollout. Event cross-attention retains the global timeline while cached visual history uses window-relative positions. We train the causal generator through teacher-guided initialization and Self-Forcing. In 60-second rollouts from 10-second training sequences, our model leads the compared systems in event-level visual quality and InternVideo2 alignment, while LongLive leads full-video Temporal-Following Score. Our code will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.