RPO-MAS: RECEIVER PREFERENCE OPTIMIZATION FOR SELECTIVE MULTI-AGENT COMMUNICATION
Abstract
Communication is a critical component of multi-agent reasoning, yet peer communication does not always improve the receiver's decision. A peer response that is correct in isolation may be redundant or harmful for a receiver with a different reasoning state, while autonomous refinement or preserving the current response may lead to better outcomes. Existing approaches for selective multi-agent communication mainly rely on source reliability, peer selection, or communication structures, providing limited supervision on whether a particular intervention benefits the receiver. In this work, we propose RPO-MAS, a receiver-conditioned preference optimization framework for selective multi-agent communication that learns intervention preferences from receiver-side outcomes. RPO-MAS formulates communication selection over three intervention types: peer-conditioned updating (SEND), autonomous refinement (SELF), and state preservation (KEEP). Starting from the same receiver state, it compares alternative interventions and constructs outcome-based preferences from their observed correctness outcomes. Based on these preferences, RPO-MAS learns a receiver-conditioned scorer to select appropriate interventions among candidate actions. Evaluation across seven reasoning benchmarks with Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct agents shows that RPO-MAS consistently improves multi-agent reasoning performance over independent aggregation and communication-based baselines. Compared with Majority Voting, RPO-MAS improves average accuracy from 76.91% to 77.72% and from 72.42% to 73.13% on the two backbones, respectively. Further analysis shows that receiver-conditioned supervision improves correction of erroneous states while reducing harmful changes to correct states.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.