MechLeash-FL: Mechanistic Leashing of Latent Features for Federated Learning under Feature Skew
Abstract
Federated Averaging (FedAvg) aggregates client models by averaging their parameters directly, so its performance implicitly depends on clients representing matching features along the same directions in the model's hidden space. We analyze what happens to this match under feature skew and representational superposition. This analysis reveals two ways that parameter averaging can degrade features: reducing feature strength and causing feature interference between clients. To mitigate these issues, we introduce MechLeash-FL, a two-phase federated training procedure that leashes clients' latent feature directions to a shared mechanistic reference during local training. MechLeash-FL uses a sparse autoencoder (SAE), shared across clients, to penalize changes in the clients' latent representations. After each aggregation, the SAE is refit to the merged model's activations and shared with the clients, providing a common reference for latent feature directions without changing the deployed model's architecture. We evaluate MechLeash-FL in two settings: a controlled toy model of superposition, where we directly measure feature shrinkage and interference, and real-data reconstruction on Office-Caltech10, a naturally feature-skewed benchmark. In a synthetic setting, MechLeash-FL reduces endpoint reconstruction error by , preserves more feature strength after aggregation, and reduces interference between clients. On Office-Caltech10, adding MechLeash-FL to FedAvg, FedProx, MOON, and Ditto reduces validation mean squared error in uncorrupted reconstruction by -.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.