Does Reward Select Interaction? Identifying and Steering Information Pathways in Multi-Agent Policies
Abstract
Multi-robot remote estimation requires allocating each robot's limited transmission quota between cached measurements and source-discovery notices. We propose a local transmission scheduler with a fixed sampler–relay covering rule. An additive score combines proxies for Gauss–Markov uncertainty and delivery probability with timeliness slack; a threshold defers low-score data. The cross-entropy method optimizes the score weights and threshold offline to minimize estimation mean squared error. A finite-horizon analysis of an isolated delivery motivates the local features and the ranking-and-threshold structure. On 200 held-out instances with 12 sources, 30 robots, and a 330-frame team budget, the scheduler reduces mean squared error by 0.0162 (95% confidence interval: [0.0089, 0.0235]), increases timely information coverage by 3.49 percentage points, and reduces mean transmissions from 326.4 to 310.6 frames relative to covering without scheduling. Independent retraining at a 360-frame budget reproduces the scheduling benefit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.