Returning Characters Need Their Original Noise in Autoregressive Video Generation
Abstract
Autoregressive video generators such as NOVA produce a long video one latent frame at a time: causal attention reads earlier frames from a key–value cache, and a flow-matching sampler turns Gaussian noise into the new frame. When a character leaves the scene and later comes back, these generators often return a plausible but different person. We trace this failure to the noise. The cached features describe the earlier frames only partly, and the details they leave open were settled by the noise that generated those frames. Fresh, independent noise redraws these details, so even an exact conditional sampler keeps an expected appearance gap of twice the variance that the conditioning leaves open. We therefore store each generated frame’s noise as an extra value in the attention cache and read it back with the attention weights the model already computes: the query that retrieves the character’s earlier features also retrieves the noise that drew them. Mixing this noise with fresh noise lets the scene keep changing, and causal flow-matching fine-tuning adapts the generator to it. On 75 leave-and-return scripts with matched NOVA models, the method, SourceValue, raises judged same-person returns from 50.2% to 75.4% at an unchanged return rate; permuting which stored noise each query reads, at matched norm, removes 15.2 points. On Wan2.1-14B, same-person returns rise from 62.40% to 83.64%, 4.00 points above the strongest streaming comparison, with 0.5% (NOVA) and 1.67% (Wan) more generation time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.