acceptodds
Under review as a conference paper at ICLR 2027

An Anchored Policy Regularizer for Deep Constrained Multi-Agent Reinforcement Learning under Tight Cost Budgets

Abstract

Reinforcement learning teams must often keep costs like energy or collisions within budget. Standard training gives each budget a Lagrange multiplier that rises under violation and falls otherwise. At binding budgets (below what an unconstrained policy spends), policy and multiplier can cycle and deep training can collapse. It is unknown which cheap stabilizer prevents this, or whether the policy survives an adversary. We propose an anchored policy regularizer that pulls each policy toward a slow copy of itself. In the tabular robust constrained Markov game framework, the existence theorem needs a weaker equilibrium notion; we prove that the one behind our multiplier update exists under robust Slater. On cooperative navigation at tight budgets, unanchored training with our multiplier update collapses after a warm-up holding the multipliers at zero, and trains without it. The anchor removes the collapse and lifts the goal rate above every constrained variant without it, with or without warm-up. Under a bounded action-perturbing adversary retrained against each final policy, anchored variants keep most of their goal rate. Training against such an adversary adds no gain beyond seed noise at that perturbation size. The method meets every budget pair of a navigation grid without per-pair tuning, in seed-pooled mean cost and per agent in two thirds of runs. We enforce the budgets on the deployed deterministic policy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.