Learning to Elicit: Co-Evolving Routing and Information Disclosure in Societies of Language Agents
Abstract
Multi-agent language models increasingly operate in societies of agents characterized by information asymmetry, where the information required to solve a task is distributed across agents. Effective coordination therefore depends not only on solving the task but also on deciding what information to elicit and from whom. We study this problem in a multi-agent meeting environment where agents possess private facts and must communicate over a shared public record to reach a verifiable settlement. We introduce ElicitBench, a benchmark in which agents hold distributed private information and must reveal the relevant facts through interaction before a designated decision maker can reach the correct outcome. We then propose Co-Adapt, a co-evolutionary training framework that jointly learns routing and agent information disclosure through reinforcement learning and privileged on-policy distillation. The router learns which agents to engage, while agents learn which private information to reveal. On the held-out test set, Co-Adapt raises the fraction of settlement checks passed from 0.65 to 0.81 and the strict success rate from 14% to 25%. Gains are largest on domains unseen during training, where the fraction of settlement checks passed increases from 0.57 to 0.78. Ablation studies further demonstrate that the gains arise from jointly learning information routing and information disclosure, rather than from either component alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.