acceptodds
Under review as a conference paper at ICLR 2027

KineTra: Fast Video Diffusion by Generating the Weights of Latent-Frame Transitions

Abstract

Video diffusion transformers generate every latent frame as a full grid of spatial tokens, so their cost grows with the number of frames times the frame size, although consecutive frames share most of their content. We propose KineTra, in which the diffusion transformer generates how latent frames change and a small network carries their content forward. KineTra splits the latent video into short segments. For each segment, the transformer generates 256 tokens that form the channel-mixing matrices of a shared executor network; the executor applies these matrices to the last latent frame of the previous segment and returns the latent frames of the current segment, whose last frame starts the next segment. The generated weights therefore encode motion and newly visible content, while appearance travels with the carried frame. A slot encoder that maps videos to these weights is trained with the executor through their own predicted boundary frames, and the transformer then learns to generate the weights with flow matching. A rank bound shows that the number of generated values must grow with the number of independently changing segments, and an error recursion bounds how errors in the carried boundary frame, the only tensor passed between segments, grow along the video. Relative to Dense on SANA-Video 2.0, LTX-Video, and CogVideoX, KineTra generates 129-frame videos 2.81–4.07× faster with VBench Quality changes of −0.05 to +0.15 points, and 961-frame videos up to 7.76× faster on a single GPU.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.