acceptodds
Under review as a conference paper at ICLR 2027

Implicit Federation Bias: A Study of Emergent Generalization in Federated Learning

Abstract

Federated learning (FL) changes the optimization dynamics by dividing training across local client trajectories and periodic aggregation, yet when and how this affects the emergence of generalization is unknown. We examine this through grokking, a delayed generalization phenomenon in which models generalize only after a prolonged period of memorization, but whose dynamics have been studied almost entirely under centralized optimization, leaving their behavior under FL unclear. We systematically study this using multiple established grokking setups under FL and find that emergent generalization persists, but its timing and behavior can change substantially. Increasing the number of clients or local optimization steps often delays generalization, while alternative federated optimizers can accelerate it by up to an order of magnitude. Stronger federation effects can produce long plateaus at sub-perfect training accuracy or, in extreme cases, no clear memorization period. In both cases, training and test performance improve together only after a long delay. To explain these effects, we analyze how local functional changes survive aggregation. Local updates are dominated by client-specific features, while aggregation preferentially suppresses these components and retains a greater fraction of globally consistent features; the accumulation of this shared global signal is associated with the onset of generalization. We characterize this as an implicit federation bias over learned structure that can delay, accelerate, or qualitatively reshape emergent generalization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.