Robust Actor-Critic Learning under Distributional Uncertainty
Abstract
We study finite-horizon Markov decision processes in which the transition dynamics are only known to lie within a Wasserstein ball around a reference model. Our goal is to learn policies that remain effective under such distributional uncertainty. A central difficulty is that the worst-case transition kernel depends on the policy itself, so the usual policy-gradient expression does not fully capture the sensitivity of the robust objective. Our main contribution is a characterization of one-sided directional derivatives of the robust value function with respect to the policy parameters, obtained via the Lagrangian dual of the inner robust problem. These derivatives replace the classical policy gradient and determine both the policy update and the auxiliary quantities that must be estimated. Building on this result, we develop a robust dynamic programming principle that applies jointly to the value function and its sensitivity, yielding backward recursions that a critic can learn to approximate. The resulting actor–critic algorithm uses three parametrized components, with separate networks for the policy, the robust value function, and its sensitivity. Under suitable regularity conditions, the latter is vector-valued and provides the policy gradient. On finite-state benchmarks the method recovers the exact robust dynamic programming solution, and on a linear–quadratic control task in continuous state space, it produces policies that remain reliable under shifted dynamics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.