Reshaping RoPE Phase of Sink Frames: Towards Long-horizon Consistency in Autoregressive Video Generation
Abstract
Autoregressive video generation (AVG) generates video content chunk by chunk, enabling long-horizon generation. In AVG, long-horizon consistency is still hard to maintain without full context. One typical solution is to keep a KV cache of beginning frames (named sink frames) and force the following generation to attend to the sink frames to maintain consistency. These methods usually suffer from the sink collapse issue, i.e., in certain steps, the model tends to unexpectedly replay the content of sink frames, rather than generating new continuous content. Typical previous works observe the correlation between RoPE encoding and the sink collapse, and propose to randomly disturb the original base of RoPE encoding. However, we observe that these methods cannot fully eliminate sink collapse. Moreover, the consequences are not stable due to the randomness and they may introduce artifacts. In this paper, we rethink the effect of RoPE on sink collapse and uncover two interesting key factors. Firstly, we carefully check the components of RoPE and find that only partial frequency pairs of RoPE dominate sink collapse phenomenon. Especially, the first pair (corresponding to the highest frequency component) is important but ignored by previous solutions. Secondly, we mathematically and empirically find that in nature, the query-sink correlation varies periodically resembling a cosine function. For different heads, the phases are largely aligned. Thus, at certain steps, all the heads may simultaneously pay high attention to the sink frames, causing sink collapse issue. Based on our findings, we propose to reshape the RoPE phase of sink frames to break the phase alignment of query-sink correlation across heads to mitigate sink collapse. Our operation is simple: we construct head-wise offsets and add them to the temporal RoPE index of sink frames when computing the query-sink attention. Consequently, the phases for different heads becomes different. Our method is training-free, deterministic and introduces no artifacts. Extensive experiments on long-horizon video generation show that our method substantially suppresses sink collapse, outperforming previous state-of-the-arts. Besides, our method largely maintains generation quality close to vanilla models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.