RUMOR Has It: Correcting Participation Bias in Decentralized Language Model Fine-Tuning
Abstract
Decentralized fine-tuning of language models can leave infrequent participants' data poorly represented. Motivated by prior user participation models that take into account model quality, we study a controlled setting in which a node becomes less likely to train and exchange its model as its loss rises relative to the network mean. This dependence can create a positive feedback loop: missed updates and model exchanges can keep the node's loss high, which can in turn further reduce its participation. To correct this imbalance using only local information, we reweight updates based on recent participation history, use past local loss as a reference to limit excessive amplification, and use disagreement with active neighbors' models as an additional signal for increasing the update weight. We call the resulting scheme RUMOR, Reference-guided Updates using Memory, Observed loss, and Reweighting. We evaluate a small transformer and seven pretrained models up to 12 billion parameters under both loss-adaptive participation and fixed participation probabilities. RUMOR substantially reduces the loss disparities associated with infrequent participation, with particularly clear benefits under strong feedback. Under weaker feedback and fixed participation probabilities, it remains competitive with aggressive participation reweighting while avoiding the severe loss increases observed for such baselines in the evaluated large-model runs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.