acceptodds
Under review as a conference paper at ICLR 2027

SWEEP: Achieving Global Convergence and Enabling Pareto Front Exploration for Deep Neural MORL

Abstract

Multi-objective reinforcement learning (MORL) has gained significant attention in recent years due to its wide range of applications. However, theoretical understanding on MORL algorithmic design remains in its infancy thus far, especially for the widely used deep neural network (DNN)-based actor–critic methods. This motivates us to study MORL from a theoretical perspective and develop a DNN-based actor–critic approach that (i) offers global convergence guarantees to weak Pareto-optimal policies and (ii) enables systematic exploration of the entire Pareto front (PF). Inspired by recent advances in multi-objective optimization (MOO) theory, our basic idea to enable systematic PF exploration is to scalarize the original vector-valued MORL problem with the weighted-Chebyshev (WC) technique and leverage the one-to-one correspondence between the PF and the WC-scalarization. However, the non-smoothness incurred by the WC-scalarization introduces difficulty in multi-objective policy gradient evaluation in MORL. Toward this end, we propose a new parameterized smoothed-WC actor-critic technique to overcome this aforementioned challenge to solve the deep-neural MORL problems with a global convergence rate of , where denotes the total number of iterations. Notably, to establish the global convergence rate guarantee for DNN-based MORL, we propose a new set of theoretical assumptions and analysis. To our knowledge, these assumptions and analysis are new in the literature and could be of independent interest to the field of actor-critic MORL with nonlinear approximation. Finally, extensive numerical experiments further validate the effectiveness of our method.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.