acceptodds
Under review as a conference paper at ICLR 2027

Latent Bridge: Asking for Help Without Words

Abstract

Text-based model collaboration requires a large model to generate intermediate responses before a smaller model can use its guidance, adding large-model decoding cost to inference. We introduce POND (Prefill-Only, No-Decode), a method that uses a frozen large sender for prefill and a small receiver for all autoregressive decoding. A trainable bridge maps the sender’s hidden states to additional attention keys and values, allowing the receiver to use these representations throughout generation. POND also supports periodic review, in which the sender processes the receiver’s partial response and updates its guidance through additional prefill passes while preserving the receiver’s existing KV cache. Across reasoning and coding benchmarks, POND improves receiver performance with senders from different model families. In particular, GLM-5.3 provides gains comparable to those of Qwen3.8-Max when paired with a Qwen receiver, supporting the effectiveness of representation transfer between independently trained model families. POND further demonstrates its ability to generalize across domains. When trained on math and coding tasks, it transfers directly to general knowledge tasks without additional domain-specific training. These results demonstrate that large frontier model prefill representations can improve small-model generation without requiring large-model autoregressive decoding. Code and inference examples can be found at https://latent-1117.pages.dev.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.