Near-Optimal Reinforcement Learning with Multi-Step Transition Lookahead
Abstract
We study reinforcement learning (RL) with transition look-ahead, where the agent may observe which states would be visited upon playing any sequence of actions before deciding its course of action. Although look-ahead can substantially improve achievable performance, [1] showed that optimal planning with multi-step transition look-ahead is NP-hard. However, this hardness was established using a discount factor close to one. It was therefore unknown whether the problem remains hard for every discount factor, and whether near-optimal planning can nevertheless be performed efficiently. We resolve both questions. First, we show that for every fixed discount factor, exact planning remains NP-hard. Second, we introduce a randomized polynomial-time approximation scheme for every fixed look-ahead depth. Third, we extend our approach to account for unknown transitions. We empirically validate the soundness of our results on the wind-farm storage-control benchmark of [2], showing that our approach, optimally accounting for -step look-ahead information, offers substantially better performance than existing algorithms. [1] Corentin Pla, Hugo Richard, Marc Abeille, Nadav Merlis, Vianney Perchet : On the Hardness of Reinforcement Learning with Transition Look-Ahead [2] Chenbei Lu, Zaiwei Chen, Tongxin Li, Chenye Wu, Adam Wierman : Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.