acceptodds
Under review as a conference paper at ICLR 2027

Returning Characters Need Their Original Noise in Autoregressive Video Generation

Abstract

Autoregressive video generators such as NOVA produce a long video one latent frame at a time: causal attention reads earlier frames from a key–value cache, and a flow-matching sampler turns Gaussian noise into the new frame. When a character leaves the scene and later comes back, these generators often return a plausible but different person. We trace this failure to the noise. The cached features describe the earlier frames only partly, and the details they leave open were settled by the noise that generated those frames. Fresh, independent noise redraws these details, so even an exact conditional sampler keeps an expected appearance gap of twice the variance that the conditioning leaves open. We therefore store each generated frame’s noise as an extra value in the attention cache and read it back with the attention weights the model already computes: the query that retrieves the character’s earlier features also retrieves the noise that drew them. Mixing this noise with fresh noise lets the scene keep changing, and causal flow-matching fine-tuning adapts the generator to it. On 75 leave-and-return scripts with matched NOVA models, the method, SourceValue, raises judged same-person returns from 50.2% to 75.4% at an unchanged return rate; permuting which stored noise each query reads, at matched norm, removes 15.2 points. On Wan2.1-14B, same-person returns rise from 62.40% to 83.64%, 4.00 points above the strongest streaming comparison, with 0.5% (NOVA) and 1.67% (Wan) more generation time.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.