acceptodds
Under review as a conference paper at ICLR 2027

Follow the Route, Know Where to Stop: Efficient Attention for Video Diffusion

Abstract

Video diffusion transformers denoise all frames of a clip together, and at every step each token attends to every other frame, at a cost quadratic in the number of frames. Efficient attention methods cut this cost for distant frame pairs and decide which pairs to keep, or how to approximate them, by comparing the two frames directly. We show that the value of a distant connection is set in the frames between: by the route the content takes, how far that route carries and where it breaks. When a sparse method keeps fewer connections, point-correspondence accuracy stays level while a quarter of the annotated point tracks can no longer be followed and motion binding drops; when propagation cannot stop, motion grows larger and temporal coherence falls. We propose PathSpan, which builds every distant connection from adjacent-frame matches: each token scores an window of the neighboring frame together with a learned no-match entry that ends propagation, two linear-time scans chain these matches, a per-channel retention sets how far each channel reads, and attention within the frame selects from the carried content. On Wan2.2, LTX-Video and Cosmos-Predict2.5, complete generation is 1.52–4.03× faster than matched dense fine-tuning with VBench totals within 0.05 points, and removing the route, the per-channel reach or the stop lowers motion binding and temporal coherence in all three generators.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.