acceptodds
Under review as a conference paper at ICLR 2027

Revise by Appending: Correcting Committed Chunks in Autoregressive Video Diffusion

Abstract

Chunk-wise autoregressive video diffusion models generate a video a few latent frames at a time and keep the key–value (KV) states of every finished chunk for the chunks that follow. Each chunk is therefore final once it is written. When a later chunk shows that an earlier one set up an event wrongly, the model can only continue from that mistake, and prompts that describe several consecutive events often end incomplete. Overwriting the earlier latent would invalidate every KV state computed from it. We observe that only the latent the Transformer attends to has to stay fixed, while the frames shown to the viewer can be decoded from the whole generated record. AmendAR uses this separation to correct earlier chunks by appending. After each video chunk, the same Transformer samples a small revision latent that describes how that exact chunk should change. The revision’s KV states are appended to the cache, so later chunks are generated in agreement with the correction, and a lightweight decoder scans the record from the last chunk to the first, adding a residual correction to each earlier latent before VAE decoding. Revisions are learned from enhanced versions of the model’s own drafts. We prove that finalizing a chunk early discards exactly the information that later chunks carry about it, and that the size of the decoder state bounds how many correction directions reach earlier chunks. On Wan2.1, MAGI-1, and HunyuanVideo 1.5 backbones, AmendAR improves event completion on VideoPhy-2, TC-Bench, and StoryEval by 4.8–8.1 points, including StoryEval from 24.9% to 33.0% on Wan, with 9–21% additional latency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.