acceptodds
Under review as a conference paper at ICLR 2027

NativeCodec: Encoding Every Video Frame with Variable-Length Causal Visual Tokens

Abstract

Reasoning about motion and object identity often depends on brief interactions that sparse frame sampling can miss. Increasing temporal resolution preserves more of this evidence, but encoding each frame independently introduces substantial visual redundancy into multimodal large language models. The central challenge is to retain temporal detail without repeatedly allocating tokens to content already represented by earlier frames. Because the amount of new information varies across frames, a compact video representation should adapt its capacity to each frame's contribution relative to its history. **We introduce NativeCodec**, a causal visual encoder for video inputs at up to 16 fps. Inspired by predictive coding, it combines complete anchor frames with variable-length incremental representations of subsequent frames. Pixel and feature reconstruction at multiple truncation lengths encourages the encoder to concentrate information complementary to its history into retained token suffixes. An offline reconstruction-based teacher supervises a per-frame length predictor, and cumulative suffix lengths determine anchor placement. **We also introduce BallBench**, which tests identity tracking through collisions between visually identical objects. With a 4B base model and a fixed suffix ratio, NativeCodec reaches 56.4 mIoU on TimeLens-Bench, 25.0 mIoU on MomentSeeker, and 26.1% accuracy on BallBench, and with predicted suffix lengths it reaches 48.7% accuracy on Video-Holmes. Relative to its Qwen3.5-4B base model, it **gains 13.9 points on BallBench, 7.6 on Video-Holmes, 6.5 on MomentSeeker, and 5.7 mIoU on TimeLens-Bench**. NativeCodec is ahead of the native pathway of its base model at every budget, and the margin widens as the budget tightens, reaching 14.2 mIoU at 4k tokens. At a suffix ratio of 1/16, NativeCodec obtains 54.4 mIoU from 16k visual tokens, while the native pathway of its base model obtains 50.9 from 64k tokens, four times the budget. The code and model will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.