Finite-Time Pareto Convergence of Decentralized Multi-Policy Gradient with Periodic Reuse for Multi-Agent Multi-Objective Reinforcement Learning
Abstract
Decentralized multi-agent policy gradient methods enable collaborative decision optimization across multiple agents without a central coordinator. However, most existing theoretical studies on decentralized policy gradients are restricted to single-objective settings, while real-world complex tasks typically require each agent to simultaneously optimize multiple conflicting objectives. To the best of our knowledge, the theoretical foundations of decentralized multi-objective policy gradient algorithms remain largely underexplored. To tackle this problem, we propose D-PRMPG, a decentralized periodic-reuse multi-policy gradient algorithm. By exploiting the reusability of dynamic multi-objective preference weights, our method substantially reduces the excessive computational overhead caused by per-iteration dynamic weight computation. We further present a rigorous convergence analysis, demonstrating that D-PRMPG attains an convergence rate under both general convex and nonconvex objective functions, where denotes the total number of iterations. Finally, empirical evaluations in multi-agent multi-objective reinforcement learning environments validate the superiority of the proposed algorithm.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.