acceptodds
Under review as a conference paper at ICLR 2027

When the Safety Signal Lies: Corrupted Cost Feedback in Constrained Multi-Agent RL

Abstract

Safe multi-agent reinforcement learning assumes that the cost signal driving constraint enforcement is reported faithfully, yet decentralized deployments receive it over a network as cost reports, an attack surface. We first show that for any connected doubly stochastic mixing matrix, an unclamped fixed point of consensus dual ascent constrains only the network average of the agents' cost estimates, not each agent's. We then establish corruption mass conservation: along a fixed primal sequence, the unprojected consensus recursion preserves the aggregate dual bias that corrupted reports inject, so topology can only redistribute it. A biased multiplier scales the constraint gradient, whereas a reward perturbation that shifts every return equally leaves the policy gradient unchanged. We propose Robust Constraint Estimation (RCE), which trims each agent's cost reports and adds a dispersion margin; under an honest-source majority it bounds the estimation error and, unlike gating updates on the untrimmed mean, cannot be stalled by over-reporting. In one instrumented testbed, persistent under-reporting from one of three sources raises true network-average cost by against a budget of (paired CI ) while the reported cost falls; a matched zero-mean perturbation does not reproduce the effect, an unconstrained control shows that the constraint holds cost near budget, and the effect increases across the tested doses . Within each block RCE removes of the effect at , where it reduces to the median, and at , where its margin is live; the margin adds clean-run conservatism but no measurable suppression, and residual harm remains. Finally, corrupting a single agent's reports, the case the spreading result addresses, yields the intended concentrated perturbation, but the realized multipliers do not follow the fixed-primal prediction: displacement is not localized on the attacked agent, and a median-deviation monitor pre-declared to catch the attack without communication detects nothing under either topology. Empirical claims are descriptive: one testbed, one operating point, three seeds.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.