acceptodds
Under review as a conference paper at ICLR 2027

The Vocabulary Bottleneck as an Estimator: Rate, Bias, and Hidden Transfer Between Language Models

Abstract

Multi-agent systems are beginning to exchange internal state instead of text to avoid both the latency and the information loss of autoregressive decoding. This creates an apparent trade-off between efficiency and safety, since latent state cannot be monitored or controlled the way text can. But the safety cost extends beyond readability. The discrete vocabulary also limits the bandwidth available for hidden coordination between agents and forces every training signal through a decoding rule that may faithfully or unfaithfully represent the teacher. To quantify what losing these constraints costs, we model the projection to vocabulary as an estimator of the sender's distribution, characterized by two properties: bandwidth per exchange and fidelity of the training signal. We test this framework at inference, where a single cache exchange replaces a decoded message, and at training, where the decoding rule is applied across a distillation corpus. At inference, a hidden signal planted in the sender's cache is unrecoverable from emitted text but recoverable at 182 bits through a trained cross-model bridge, and every projection-based narrowing we test retains most of that capacity at two to three times the cost of not communicating. At training, no choice of decoding rule blocks hidden trait transfer. Sampling recovers the teacher's full distribution over many examples, consistent with prior findings that transfer survives filtering and paraphrasing. Argmax decoding produces a systematically biased target that displaces the student from its teacher by factors of 51.9 to 689.7 across three architectures, a phenomenon we call mode drift, while relocating the acquired disposition rather than suppressing it. Soft-target distillation avoids this displacement but tolerates far more training before degeneration, at 3.6% collapse versus 91.7% for argmax at matched budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.