Networked Restless Multi-Armed Bandits with Delayed-Diffused Externality
Abstract
We study networked restless multi-armed bandits in which actions generate externalities that are delayed in time and diffuse over a graph. We formalize this setting as a Delayed-Diffused Externality Restless Multi-Armed Bandit (DDE-RMAB), where activating an arm injects a latent exposure signal that persists, propagates through the network, and modulates future arm dynamics. This state-level coupling breaks the single-arm decomposition underlying classical Whittle-style methods. To address this challenge, we propose the Propagation-Adjusted Index (PAI), a scalable index policy that augments a local activation advantage with an analytic correction for long-run network spillovers. The correction is derived as a graph resolvent applied to local exposure marginal values, yielding an interpretable decomposition into private local value and delayed propagation value. We provide an approximation analysis showing that, under stable propagation and regular local modulation, the performance gap of PAI is controlled by explicit errors from value decomposition, multi-action interaction, propagation truncation, and marginal-value estimation. Experiments on positive and negative externality environments, as well as a semi-real graph benchmark, show that PAI is consistently best or near-best among the state-of-the-art baselines. These results highlight the importance of explicitly modeling delayed graph diffusion in scalable control of networked restless multi-armed bandits.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.