The Vocabulary Bottleneck as an Estimator: Rate, Bias, and Hidden Transfer Between Language Models
Abstract
Multi-agent systems are beginning to exchange internal state instead of text to avoid both the latency and the information loss of autoregressive decoding. This creates an apparent trade-off between efficiency and safety, since latent state cannot be monitored or controlled the way text can. But the safety cost extends beyond readability. The discrete vocabulary also limits the bandwidth available for hidden coordination between agents and forces every training signal through a decoding rule that may faithfully or unfaithfully represent the teacher. To quantify what losing these constraints costs, we model the projection to vocabulary as an estimator of the sender's distribution, characterized by two properties: bandwidth per exchange and fidelity of the training signal. We test this framework at inference, where a single cache exchange replaces a decoded message, and at training, where the decoding rule is applied across a distillation corpus. At inference, a hidden signal planted in the sender's cache is unrecoverable from emitted text but recoverable at 182 bits through a trained cross-model bridge, and every projection-based narrowing we test retains most of that capacity at two to three times the cost of not communicating. At training, no choice of decoding rule blocks hidden trait transfer. Sampling recovers the teacher's full distribution over many examples, consistent with prior findings that transfer survives filtering and paraphrasing. Argmax decoding produces a systematically biased target that displaces the student from its teacher by factors of 51.9 to 689.7 across three architectures, a phenomenon we call mode drift, while relocating the acquired disposition rather than suppressing it. Soft-target distillation avoids this displacement but tolerates far more training before degeneration, at 3.6% collapse versus 91.7% for argmax at matched budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.