FedArena: A Unified Agent Harness for Federated Learning Poisoning Research
Abstract
Evaluating robust aggregation in federated learning (FL) requires attacks that adapt to the defense across clients and training rounds. LLM agents offer a way to explore such adaptations, but interpreting their results requires knowing what they observed, what they changed, and which changes actually affected training. We introduce FedArena, a unified harness for controlled, auditable agent-driven poisoning research. FedArena represents attack and defense operations as typed, composable objects, filters observations and interventions by an explicit threat model, and separates local defensive trials from real training rounds. It supports two agent modes, revising attack programs between complete FL runs (L1) and editing client messages between real rounds (L2), and compares both with scripted and human-designed attacks (L0). In an L1 case on CIFAR-10 against Divide-and-Conquer (DnC), a spectral outlier-filtering defense, the agent adds cross-round memory to an attack program; disabling the memory raises development accuracy by 1.98 pp. The frozen program, which also queries a loss-gradient oracle, ends with lower accuracy than the Min-Max attack on 4 of 5 seeds (54.5% vs. 58.0%). In 35 GPT-5.5 episodes on MNIST under Krum, adding local defensive trials raises the rate at which Krum admits controlled updates from 67% to 89%, yet mean final accuracy drops stay at or below 1.34 pp, below a scripted fixed program (3.74 pp) and Min-Max (7.89 pp). Statistics and full-message views reduce token use by 43–45% without increasing damage. Admission and damage thus diverge, and agent-driven adaptive evaluation should report both.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.