Swiftly Generating Whittle Index in a Changing Environment: An Environment Parameter-Driven Actor-Critic Approach
Abstract
Whittle index provides a low-complexity and adaptive solution to large-scale restless multi-armed bandit (RMAB) problems. However, in practice, when the environment faced by each arm constantly changes, the Whittle index has to be repeatedly recomputed, which incurs exceedingly high computational complexity. In this paper, we develop a new environment parameter-driven neural-network approach for generating the Whittle index for constantly-changing environments. Our neural network only needs to be trained once, and it can instantly produce the Whittle index for a new environment in real-time, as soon as the new environment can be identified by its environment parameter (which can be estimated by either a hand-crafted estimator or a trained environment encoder). We further provide a new Whittle Actor-Critic (WAC) pipeline to train the index network. Compared to the prior Whittle Q-Learning pipeline, our new WAC pipeline is less sensitive to the neural-network size and can produce meaningful index-like policies even for non-indexable problems. Numerical results show that our WAC outperforms other baselines in a broad range of RMAB settings in terms of both index accuracy and control performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.