Self-Play Enhancement via Advantage-Weighted Refinement in Online Federated LLM Fine-Tuning
Abstract
Recent works have advanced feedback-based learning systems, whereby a foundation model is able to intake incoming feedback (e.g., from a user) to self-improve, creating a self-loop system of training. However, existing works are limited in needing to consider an offline setup to allow for such feedback-based methods, requiring privileged ground-truth contexts for training. Moreover, there is limited consideration of federated learning (FL), which is particularly well-suited for incorporating external feedback across large networks of end users, for example, but requires methods to be efficient for training on resource-constrained edge devices. Therefore, we introduce SPEAR (Self-Play Enhancement via Advantage-Weighted Refinement), an efficient online learning algorithm for federated LLM fine-tuning. SPEAR utilizes a feedback-guided self-play loop to construct naturally contrastive pairs per prompt which are used to train the model with (i) standard maximum likelihood on satisfactory completions and (ii) confidence-weighted unlikelihood on tail tokens of unsatisfactory completions. SPEAR requires only an interaction-level binary satisfaction signal together with non-answer feedback, without requiring a ground-truth answer to be provided, and does not need expensive group generations, meaning training both online and in a resource-efficient manner is feasible. We validate SPEAR across various benchmark datasets, demonstrating its superior performance in comparison to relevant baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.