Outlier Capping for One-shot Federated LoRA: Dominated, Not Misaligned
Abstract
One-shot federated merging combines LoRA adapters that clients trained independently into one adapter, in a single round and without data, gradients, or a validation set. In this setting recent merge operators collapse, and on eight financial adapters for Llama-3.1-8B and hub adapters for Mistral-7B several of them score exactly zero. Realigning client subspaces is a natural remedy, but we show that misalignment is not the cause. The cause is client dominance. Update norms differ by on average, a few clients carry nearly all the energy of the summed update, and scaling a single client reproduces the collapse. A controlled training study further shows that which LoRA factor looks shared across clients depends on the initialization protocol, so a fix should act on the product . We propose \oursfull (\ours), a data-free server-side step for one-shot federated LoRA that shrinks only the clients whose update norm is an outlier of the pool. It bounds dominance but does not remove it. The alternative bounds (clipping at the median as in robust federated learning, shrinking the whole update, and equalizing every client) fail in the ways the diagnosis predicts. All operators we test except task arithmetic and DC-Merge (which flattens itself) also need each client's spectrum to be flattened first (PolCap). With task arithmetic, has the highest retention among the rules that regress on neither pool, and the best worst-case regression against the unmerged base among the rules we measure on every axis at the largest draw.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.