CoVisco: Codec-Native Vision Encoder with Native Token Compression for Efficient Unified Visual Understanding
Abstract
Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with frames times patches per frame, and long-video understanding quickly becomes computationally prohibitive. Existing remedies compress visual tokens after the vision encoder and before the language model—an inherently lossy post-hoc pruning of a representation that was never trained to be compressed, with particularly severe information loss on video. Recent work has moved compression into encoder pretraining, but does so by discarding most fine-grained tokens outright, conflating compression with deletion. We present \model, a unified image-video codec-native vision encoder that performs native token compression: during pretraining, each video segment is equipped with a small set of learnable abstract tokens that learn to absorb the segment's semantics, while all fine-grained patch tokens retain their full information flow—compression and fine-grained preservation coexist rather than compete. Cross-segment interaction flows exclusively through abstract tokens, providing a long-video modeling scheme that avoids dense cross-frame full attention. For MLLM deployment, a lightweight token selector further enables a family of token strategies that combine abstract tokens with top- fine-grained tokens selected on demand, dynamically trading accuracy for the visual-token budget. In codec mode, the selector operates on the tokens already retained by codec-based input selection, yielding a two-stage compression pipeline before the LLM: codec-based selection first reduces the visual input, while the learned abstract-token interface and selector further reduce the visual sequence exposed to the LLM. For uniformly sampled and frame-collage inputs, the same learned second stage operates directly on the encoder's fine-grained token stream. Together, these components provide multi-level visual modeling: patch tokens preserve fine-grained evidence, abstract tokens encode segment-level semantics, and abstract-mediated interaction builds video-level context. Pretrained with three contrastive objectives on 565M image–text pairs, and 6.4M videos, is competitive with selected larger embedding models on video-oriented evaluation while using a compact vision tower. It retains useful image-side retrieval and classification performance, and compresses 64-frame videos into as few as 400 visual tokens (2.4% retention) while achieving video understanding performance comparable to, and in some cases better than, OneVision-Encoder under the evaluated settings. These results highlight its ability to support efficient unified visual understanding across both images and videos.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.