acceptodds
Under review as a conference paper at ICLR 2027

Towards End-to-End Pretraining of Streaming Video Models from Sparse Causal Histories

Abstract

We propose Sparse Predictive Transformers (SPT), a simple framework for efficient, end-to-end pretraining of streaming video generators and video representations. A single flow loss jointly trains a ViT, a Causal Transformer, and a DiT to predict a complete chunk sampled from a short future horizon using sparse history features. Streaming generation reuses the history state across denoising steps. At 355M and 775M parameters, SPT achieves lower FVD than teacher-forcing and diffusion-forcing baselines with roughly half their training compute and lower streaming latency. Generative pretraining also improves frozen history-encoder SSv2 linear-probe Top-1 by 12.71 points over random initialization with the same pretrained codec. Pixel-space rendered-text generation and GLUE transfer provide initial evidence for language modeling through video.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.