BLoM: Bayesian Inference over Merge Operators for LoRA Model Merging
Abstract
Merging low-rank adaptation (LoRA) adapters efficiently combines task-specific capabilities, but common practice selects a discrete merge configuration specifying coefficients, calibration strengths, and ranks, using validation accuracy. Such selection can overfit when task updates interfere. We quantify this merge-selection gap and observe losses up to 5.6 accuracy points on a high-interference task pair. BLoM replaces selection among discrete configurations with posterior inference over the continuous merge operator that each configuration induces. A lossless core-space projection makes this operator compact and enables a full-covariance Gaussian approximation within a representative layer of models with up to 8 billion parameters. The posterior supports a regularized mean predictor, Bayesian model averaging (BMA), confidence-based abstention, and PAC-Bayes analysis of Gibbs risk, the expected loss of a posterior-sampled operator. Across five standard settings spanning three language models and two vision transformers, BLoM-Mean improves the NLI suite average by 6.9 points over the strongest of twelve baselines, including task arithmetic, TIES, DARE-TIES, RegMean, aligned-subspace, and LoRA-specific methods. In cross-domain merging, it gains 13.9 points over the strongest of three evaluated baselines, task arithmetic, TIES, and TSV-M. On the pilot pair, BLoM-Mean improves validation-selected TSV-M by 4.8 points and exceeds its finite-grid test oracle. Nominal balanced-design risk curves remain below one across nine verified ranks. These results recast LoRA merging as inference over plausible merge operators.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.