acceptodds
Under review as a conference paper at ICLR 2027

AMPO: Adaptive Minimax Policy Optimization for Federated Reinforcement Learning

Abstract

In federated reinforcement learning (FRL), we aim to learn a shared policy that performs reliably across all local environments. We therefore focus on the worst-case return, rather than the average return optimized by conventional methods. Achieving this is challenging since the worst-case environment is induced by the learned policy itself and shifts during training. We formulate this goal as an adaptive minimax objective and address the resulting non-stationarity through a saddle-point reformulation: we propose adaptive minimax policy optimization (AMPO), a primal–dual policy gradient method in which the dual variable tracks the shifting worst-case environment. Our analysis both guarantees convergence—to a stationary point with sublinear dual regret—and informs the algorithm design: it prescribes how fast the dual variable may be updated relative to the policy, and quantifies a robustness–scalability trade-off in which emphasizing difficult environments reduces the effective parallelism gained from federation. Experiments on heterogeneous MuJoCo benchmarks show that AMPO improves worst-case performance over average-return baselines while retaining competitive average performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.