acceptodds
Under review as a conference paper at ICLR 2027

CausalFold: Letting Later Video Blocks Read and Train Earlier Attention

Abstract

Autoregressive video diffusion models generate a video one latent block at a time, and each new block attends to earlier blocks through a key–value (KV) cache. In every attention layer, the cache returns the Values of earlier tokens, which are projections of their input to that layer. The output that an earlier token computes in the same layer, for example the corridor content that a doorway token selects with its own Query, reaches later blocks only through the next layer, and a later block’s loss reaches that attention only through deeper layers. We introduce CausalFold, which keeps each cached Key and stores the token’s attention output in place of its Value. A token then reads the Values of its own block, whose tokens are denoised jointly, and the outputs of earlier blocks, which are already final. The result is a recursion over blocks that ends after finitely many reads, without truncation or discounting. Each output remains a convex combination of Values, and the Jacobian of a later output with respect to an earlier token’s attention logits is the probability that attention paths visit that token times the differences among its candidates. CausalFold adds no parameters and keeps each backbone’s masks, sampler, and cache schedule. Against Standard-KV controls with identical training on five backbones, it raises transition completion by 2.3 to 4.9 points in all six text- and image-to-video comparisons, including from 37.9 to 42.7 for text-to-video on Wan2.1-1.3B at 5.4% longer generation time. In image-to-video, the trained model depends on the cached outputs, and removing the later-block gradient on earlier attention removes about half of the gain.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.