From Influence to Utility: Reference-Aligned Message Scheduling for Frozen Multi-Agent Policies
Abstract
A learned communication policy can generate more messages than a deployment link can carry. Selecting among them requires distinguishing a useful receiver response from a merely large one. We study this problem with the actor and communication stack frozen. CA-RAPU evaluates each candidate-induced policy shift against an action reference from other senders addressing the same receiver, then combines positive progress with influence, delivery reliability, confidence, and cost. The analysis characterizes the ambiguity of unsigned influence and gives sufficient conditions under which reference error preserves the selected set. Across 20 independently trained stacks in an assignment-controlled Spread task, CA-RAPU improves return over unsigned KL by 0.0277 [0.0153, 0.0401] at identical traffic. When each method calibrates its own transmission threshold, the net-return gain is 0.132 [0.112, 0.153]. Brackets give paired 95% confidence intervals. Redistributing peer support reverses the comparison with KL even when queue size stays fixed. Reference alignment thus helps allocate a frozen policy's messages when peers provide informative evidence about receiver actions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.