acceptodds
Under review as a conference paper at ICLR 2027

Beyond Single-Method Merging: Hierarchical Chains and Capability-Guarded Evaluation of LoRA Experts

Abstract

Model merging offers a nearly training-free alternative to costly joint fine-tuning, yet practical guidance on which merging recipe to use remains anecdotal. Furthermore, merged models can quietly trade away general capabilities for target-domain gains. We conduct an exhaustive, guard-railed registry sweep: all twenty merge methods exposed by the mergekit toolkit are natively exercised against four real LoRA experts (Chinese, math, social, science) fine-tuned on real TMMLU+ data from gemma-4-26B-A4B-it. Across 105 structural configurations (singles, hierarchies, sequential chains, and a passthrough control), each recipe is scored on TMMLU+ under a strict three-metric capability contract (MMLU, GSM8K, IFEval). A subsequent refinement stage optimizes the best recipes via anchored weighting and layer-adaptive scaling (44 additional configurations). Empirically, our results strongly validate the efficacy of merging: the best confirmed recipe (a performance-anchored DELLA to RAM chain) elevates macro TMMLU+ from 0.629 (base) to 0.720, passes all capability gates, and improves unseen-dataset transfer (ARC: 0.969 vs. 0.945). Crucially, our rigorous audit surfaces a systemic cautionary finding: across three independent runs, small-sample screening rankings fail to predict confirmation rankings (Spearman ρ ≤ 0.30), underscoring that leaderboard evaluations of merge recipes without explicit large-sample confirmation are inherently unreliable. We release all YAML recipes, scores, and audit artifacts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.