acceptodds
Under review as a conference paper at ICLR 2027

Connectors as Channels: Usable Information between Frozen Vision and Language

Abstract

Vision language models couple a frozen visual encoder to a frozen language model through a small trainable connector. The connector family and the number of retained visual tokens are still chosen by ablation. Published comparisons disagree on which of these choices matters, and they still lack a common account of when a given design is appropriate. The information bottleneck is the classical theory of this kind of compression. It prices a learnable code, while a connector sits between two frozen spaces and is read only through the directions that move the next token loss. In this paper, we model the connector as a channel and measure usable information, the mutual information that remains after the frozen language model has read the mapped tokens. On released Qwen3-VL checkpoints, the readable set recovered from last layer hidden states is much thinner than the leading principal subspace of the embedding table. Identifying this set exposes a concrete cost of query mixing, because lengthening the encoder without new signal dilutes the attention mass on a signal set and lowers usable information. We then estimate canonical spectra from teacher forcing hidden states and from caption vectors, and we train linear, two layer, and query connectors on frozen CLIP features that feed a frozen language model. These results account for the large token counts retained by published designs, for the sufficiency of a single matrix when the encoder is already contrastively aligned, and for the behaviour of a query bottleneck as a mixing mechanism rather than an integer token budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.