Safety after Decentralization: A Theory of Policy Extraction
Abstract
Offline multi-agent reinforcement learning often learns a joint distribution of states and actions, then fits a separate policy for each agent. Acting independently changes which actions occur together and which states the team visits. As a result, a stronger safety penalty can lower the training target's cost while raising the deployed policy's cost, even when optimization and policy fitting are exact. In the one-state setting with Kullback-Leibler (KL) regularization, binary actions, and binary costs, every type of cost that allows such a reversal has an example using at most four joint actions; four are sometimes necessary. We classify the costs for which a stronger penalty never raises deployed cost. We also identify when no penalty can reach a budget that an independent policy can satisfy. These findings motivate Deployed-Cost Projection (DCP), which adjusts the joint target using the gradient of deployed cost, preserves the state-flow constraints, and refits the local policies. We derive this gradient and give conditions for cost reduction when the gradient, projection, and policy fitting are approximate. In a structured safety task with up to 64 agents, DCP reaches budgets that penalty increases cannot reach. Controlled comparisons show how cross-agent action covariance helps DCP retain more reward at the same cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.