D-Tok: Single-Pass Dynamic-Length 1D Video Diffusion Tokenizer
Abstract
Fixed-length visual tokenization with 2D/3D inductive bias is suboptimal in learning rich semantics and facilitating visual autoregressive modeling. 1D visual tokenizers have mitigated these issues; however, test-time search for optimal token length locks its dynamics and efficiency. We propose D-Tok, a single-pass ynamic-length 1 video iffusion tokenizer. The core idea is to map visual content to a set of latent tokens and then discard redundant tokens using a single-pass end-of-sequence detector. The remaining tokens are decoded via the diffusion decoder. Experimental results show that D-Tok surpasses all baselines, setting a new state-of-the-art rFVD/gFVD among video tokenizers with the highest token compression rate. It also achieves comparable performance to its fixed-length counterparts while enabling semantic representations. Downstream tasks on video generation demonstrate that our method achieves better results on autoregressive modeling compared to 2D or 3D methods. Code will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.