acceptodds
Under review as a conference paper at ICLR 2027

Robust Planning in Stochastic Environments via Selective Adversarial Interventions

Abstract

Stochastic MuZero and Stochastic AlphaZero have achieved the state-of-the-art for planning in stochastic environments by modeling stochasticity using afterstates and chance events. Their expectation-based training focuses on maximizing average outcomes, at the expense of robustness to rare but consequential chance events. To address this, we introduce Robust Stochastic Zero (RSZ), the first method to extend Stochastic AlphaZero using robust adversarial reinforcement learning (RARL) with selective adversarial interventions. Unlike typical RARL methods that constantly perturb states, RSZ intervenes only at the most critical afterstates, which rarely occur but cause significant value degradation. Furthermore, it leaves the otherwise dynamics untouched, minimizing the interventions to avoid distorting the inherent stochastic dynamics. This approach improves the planning robustness without obstructing the ability to plan in the natural environment. On benchmark stochastic environments, 2048, Tetris Block Puzzle, and Server Allocation, RSZ achieves improved performance over the baseline with intervening in only 0.01% of events, and remains comparable when no interventions are applied. Our findings demonstrate that the right rather than constant interventions are a direction for robust planning in stochastic environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.