States as Messages: Latent Communication for State Space Models
Abstract
State space models (SSMs) are increasingly used in language modeling, and when model instances hold different contexts, one may need information that another has already processed. Transferring this information through natural-language summaries requires autoregressive generation at the source and re-encoding at the receiver. Latent communication can reduce this overhead by directly exchanging internal representations such as the fixed-size recurrent states of SSMs. Using these states for communication requires the receiver to combine the source's state with its own while retaining information from both contexts. We introduce a dynamics-grounded fuser for pretrained multilayer SSMs. The fuser learns a lightweight per-channel correction to a single-layer state decomposition while keeping the base model weights frozen. Across five scales each of Mamba and Mamba-2, our method closes up to 75% and 84% of the receiver-only to full-context perplexity gap on WikiText-103, respectively. The method generalizes across sequence lengths, partition ratios, and domains without retraining. On separate-document QA with HotpotQA and MuSiQue, task-specific state communication achieves exact-match scores comparable to natural-language summaries and closes up to 74% of the output-distribution gap to the full-context teacher. On HotpotQA, state communication runs approximately faster end to end than the summary baselines on a single GPU, with no generated message tokens.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.