acceptodds
Under review as a conference paper at ICLR 2027

Finite-Time Pareto Convergence of Decentralized Multi-Policy Gradient with Periodic Reuse for Multi-Agent Multi-Objective Reinforcement Learning

Abstract

Decentralized multi-agent policy gradient methods enable collaborative decision optimization across multiple agents without a central coordinator. However, most existing theoretical studies on decentralized policy gradients are restricted to single-objective settings, while real-world complex tasks typically require each agent to simultaneously optimize multiple conflicting objectives. To the best of our knowledge, the theoretical foundations of decentralized multi-objective policy gradient algorithms remain largely underexplored. To tackle this problem, we propose D-PRMPG, a decentralized periodic-reuse multi-policy gradient algorithm. By exploiting the reusability of dynamic multi-objective preference weights, our method substantially reduces the excessive computational overhead caused by per-iteration dynamic weight computation. We further present a rigorous convergence analysis, demonstrating that D-PRMPG attains an convergence rate under both general convex and nonconvex objective functions, where denotes the total number of iterations. Finally, empirical evaluations in multi-agent multi-objective reinforcement learning environments validate the superiority of the proposed algorithm.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.