acceptodds
Under review as a conference paper at ICLR 2027

Learn and Pin: A Transferable JEPA-Style Visual Token Compressor for Frozen VLM Hosts

Abstract

In vision–language models, the image dominates the input: LLaVA-1.5 renders one image as tokens, far more than a typical prompt, so most image compute is spent inside the language model. Existing visual token compressors either hand-specify the pre-designed method or learn it with task supervision tied to one host. None learns label-free model inside a frozen encoder and transfers it, unmodified, to a new host. In this paper, we train a small compressor inside a frozen vision encoder without captions. It uses a JEPA-style predictive term and a distributional pin. The pin is a kernel two-sample statistic that keeps the distribution of the compressed tokens close to the encoder's own. The compressor then runs frozen in the encoder's forward pass and is evaluated zero-shot. We do not quote published numbers. We re-run all competitors, both published methods and training-free anchors, in one setting. They are matched on token count, or on total cost for decoder-side pruners. At (a quarter of the tokens) on the LLaVA + SmolLM2-135M parameter host, the proposed method leads the field by CIDEr zero-shot and one-shot at of the uncompressed analytic cost. The compressor transfers unchanged into LLaVA-1.5-7B. It retains – of the uncompressed ceiling on five of six in-context tasks and – on five standard benchmarks, at of the analytic cost. It also transfers to an unseen Q-Former bridge at retention, over the strongest competitor. On CLIP retrieval, it retains of recall at and leads by .

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.