acceptodds
Under review as a conference paper at ICLR 2027

Robust Large-Scale Multi-agent Reinforcement Learning against Byzantine Fault

Abstract

In a large-scale multi-agent system, several corrupted agents with unknown identities can derail the system by spreading abnormal behaviors that harm coordination. Such failures are known as Byzantine faults, under which defenders know neither the number nor the identities of agents being attacked. The classic solution concept is a Bayesian equilibrium, where defenders act on beliefs about whether each agent is cooperative or adversarial. However, this inference becomes intractable as the agent population grows, since the number of possible corrupted subsets increases exponentially. Our key insight is that individual defenders do not need to identify the type of the whole team, but only to aggregate their type uncertainty into the mean field, capturing how the population is distorted rather than by whom. Here we propose Belief over Mean Field (BeMF), which realizes this as robust mean-field control over two agent types. BeMF assigns each agent a Byzantine risk score in linear time, uses these scores to build a belief-weighted mean field, and learns a shared policy that acts on it. We prove that BeMF's belief-weighted mean-field model approximates both population behavior and worst-case performance with explicit error bounds that shrink as the population grows, so the attacker–defender equilibrium remains approximately valid at finite population sizes. We evaluate BeMF on red-versus-blue team combat with up to 144 agents and power-grid control with 20 to 40 agents. In both domains, BeMF outperforms representative baselines even when half of the agents are corrupted. These results establish BeMF as a scalable approach to robust multi-agent cooperation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.