acceptodds
Under review as a conference paper at ICLR 2027

Does Reward Select Interaction? Identifying and Steering Information Pathways in Multi-Agent Policies

Abstract

Multi-robot remote estimation requires allocating each robot's limited transmission quota between cached measurements and source-discovery notices. We propose a local transmission scheduler with a fixed sampler–relay covering rule. An additive score combines proxies for Gauss–Markov uncertainty and delivery probability with timeliness slack; a threshold defers low-score data. The cross-entropy method optimizes the score weights and threshold offline to minimize estimation mean squared error. A finite-horizon analysis of an isolated delivery motivates the local features and the ranking-and-threshold structure. On 200 held-out instances with 12 sources, 30 robots, and a 330-frame team budget, the scheduler reduces mean squared error by 0.0162 (95% confidence interval: [0.0089, 0.0235]), increases timely information coverage by 3.49 percentage points, and reduces mean transmissions from 326.4 to 310.6 frames relative to covering without scheduling. Independent retraining at a 360-frame budget reproduces the scheduling benefit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.