acceptodds
Under review as a conference paper at ICLR 2027

Swiftly Generating Whittle Index in a Changing Environment: An Environment Parameter-Driven Actor-Critic Approach

Abstract

Whittle index provides a low-complexity and adaptive solution to large-scale restless multi-armed bandit (RMAB) problems. However, in practice, when the environment faced by each arm constantly changes, the Whittle index has to be repeatedly recomputed, which incurs exceedingly high computational complexity. In this paper, we develop a new environment parameter-driven neural-network approach for generating the Whittle index for constantly-changing environments. Our neural network only needs to be trained once, and it can instantly produce the Whittle index for a new environment in real-time, as soon as the new environment can be identified by its environment parameter (which can be estimated by either a hand-crafted estimator or a trained environment encoder). We further provide a new Whittle Actor-Critic (WAC) pipeline to train the index network. Compared to the prior Whittle Q-Learning pipeline, our new WAC pipeline is less sensitive to the neural-network size and can produce meaningful index-like policies even for non-indexable problems. Numerical results show that our WAC outperforms other baselines in a broad range of RMAB settings in terms of both index accuracy and control performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.