SWIN: Seed Weight Integration for LLM Fine-Tuning
Abstract
Fine-tuning large language models exhibits high variance across random seeds, forcing repeated runs or expensive ensembles to reach top performance. We introduce SWIN (Seed Weight INtegration), a training-free, geometry-aware fusion method that integrates models fine-tuned with different seeds without increasing inference cost, motivated by the largely low-rank nature of fine-tuning updates and the emergence of complementary, task-relevant directions across seeds. SWIN first computes per-layer low-rank SVDs and forms basis-invariant projection operators to estimate a shared left/right consensus subspace. In that shared coordinate system it decouples each seed’s core update into orientation and scale via polar decomposition, aggregates orientations with a Procrustes projection and averages symmetric scales, and reconstructs a single low-rank update. By preserving matrix geometry and avoiding brittle element-wise averaging, SWIN effectively merges complementary, task-relevant directions while suppressing redundant spectral tails. Across mathematical reasoning, commonsense reasoning, code generation, and RL benchmarks, SWIN consistently matches or exceeds the best individual seed and outperforms element-wise baselines (e.g., improving MetaMath and AIME by up to +4.0 and +9.9 absolute points respectively) with no extra training or inference overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.