Post-Training with Communication
Abstract
We study how to post-train language models on tasks whose evidence no single model holds. We propose post-training with communication, a training paradigm in which identical copies of one model, the agents, each hold a different shard of the context, exchange free-text messages on a shared board, and then answer. No agent has a role, and no agent holds all the evidence, so the team earns full reward only by communicating. Training uses only the verifier of single-model RLVR, with no reward model and no hand-written collaboration reward. To further improve performance, we add two counterfactual credit terms, which the same verifier computes. We find that the trained team scales zero-shot. Trained with only four agents, it keeps most of its accuracy in teams of 32 and remains effective with 64, on inputs several times longer than the model's context window. At these scales, on long-context memory and multi-hop question answering, the team matches or outperforms the same model with a YaRN-extended context window or run as map-reduce. The team agent also stays a generalist: it has no fixed role, and communication training causes no measurable loss of general reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.