acceptodds
Under review as a conference paper at ICLR 2027

Post-Training with Communication

Abstract

We study how to post-train language models on tasks whose evidence no single model holds. We propose post-training with communication, a training paradigm in which identical copies of one model, the agents, each hold a different shard of the context, exchange free-text messages on a shared board, and then answer. No agent has a role, and no agent holds all the evidence, so the team earns full reward only by communicating. Training uses only the verifier of single-model RLVR, with no reward model and no hand-written collaboration reward. To further improve performance, we add two counterfactual credit terms, which the same verifier computes. We find that the trained team scales zero-shot. Trained with only four agents, it keeps most of its accuracy in teams of 32 and remains effective with 64, on inputs several times longer than the model's context window. At these scales, on long-context memory and multi-hop question answering, the team matches or outperforms the same model with a YaRN-extended context window or run as map-reduce. The team agent also stays a generalist: it has no fixed role, and communication training causes no measurable loss of general reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.