BACKDOOR ECHOES: SAME MODEL, DIFFERENT SECURITY FUTURES IN FEDERATED LEARNING
Abstract
Federated learning models can retain backdoors even after malicious clients leave and training continues with only benign participants. Existing studies primarily attribute such persistence to model parameters, overlooking server optimizer states that encode historical updates and influence future training. We show that server optimizer state can substantially alter post-attack backdoor behavior even when model parameters are identical. Through causal state interventions, we observe an average 8.82% attack-success-rate (ASR) gap across independently pretrained image classifiers. On Llama-2-7B, this gap reaches 85.99% after twenty benign rounds, while clean behavior remains nearly identical. Motivated by this state dependence, we develop DYNAPHANTOM, a history-aware attack that conditions malicious submissions on previous submissions to indirectly shape the hidden server state without accessing it. Our findings identify optimizer state as an overlooked source of persistent backdoor risk and show that restoring identical model weights does not necessarily restore the same security trajectory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.