Latent Bridge: Asking for Help Without Words
Abstract
Text-based model collaboration requires a large model to generate intermediate responses before a smaller model can use its guidance, adding large-model decoding cost to inference. We introduce POND (Prefill-Only, No-Decode), a method that uses a frozen large sender for prefill and a small receiver for all autoregressive decoding. A trainable bridge maps the sender’s hidden states to additional attention keys and values, allowing the receiver to use these representations throughout generation. POND also supports periodic review, in which the sender processes the receiver’s partial response and updates its guidance through additional prefill passes while preserving the receiver’s existing KV cache. Across reasoning and coding benchmarks, POND improves receiver performance with senders from different model families. In particular, GLM-5.3 provides gains comparable to those of Qwen3.8-Max when paired with a Qwen receiver, supporting the effectiveness of representation transfer between independently trained model families. POND further demonstrates its ability to generalize across domains. When trained on math and coding tasks, it transfers directly to general knowledge tasks without additional domain-specific training. These results demonstrate that large frontier model prefill representations can improve small-model generation without requiring large-model autoregressive decoding. Code and inference examples can be found at https://latent-1117.pages.dev.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.