acceptodds
Under review as a conference paper at ICLR 2027

CodecTok: Codec-informed Adaptive Token Allocation for Efficient Video Generation

Abstract

Discrete video tokenization is a central efficiency bottleneck in autoregressive video generation, as each clip must be compressed into a token sequence whose length strongly affects downstream decoding cost. Fixed-budget video tokenizers allocate the same number of tokens to every clip, which can use capacity inefficiently across videos with different spatiotemporal complexity. We draw inspiration from rate control in classical video codecs: inexpensive measures of spatial and temporal complexity can guide content-adaptive token allocation. Based on this view, we introduce CodecTok, a codec-informed adaptive video tokenizer that adapts the idea of content-dependent capacity allocation to discrete video representations. CodecTok uses a training-free allocator to estimate spatial complexity from block-wise gradient statistics and temporal complexity from Hadamard-domain inter-frame residuals, producing a content-adaptive token budget for each clip without training an auxiliary allocation network or performing per-clip budget search. To support variable budgets in practice, CodecTok further employs a padding-free flexible backbone with sequence packing and 3D rotary positional encoding, enabling arbitrary resolutions and variable-length token sequences. A semantic alignment objective distills structured representations from a frozen video foundation model, encouraging highly compressed tokens to preserve meaningful spatiotemporal semantics. Experiments on UCF-101 demonstrate competitive reconstruction and downstream AR generation at reduced token budgets. With 743 reconstruction tokens and 730 generated tokens on average, CodecTok achieves an rFVD of 16 and a gFVD of 44, supporting a favorable balance between visual quality and token efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.