acceptodds
Under review as a conference paper at ICLR 2027

Selective Modular Merging of Heterogeneous Language Models

Abstract

Model merging offers a training-free strategy for model adaptation and knowledge transfer, with a significant efficiency advantage over existing post-training methodsĀ (such as reinforcement learning-based distillation and supervised fine-tuning). However, current merging methods either apply to homogeneous models only or require additional data and training costs, limiting their feasibility and practical usefulness. In this study, we propose Selective Modular Merging (SMM), a simple yet effective training-free method for merging heterogeneous language models. In principle, given a task-specific student model and its corresponding task vector, SMM further improves performance on the task by selectively merging a heterogeneous but more powerful source model into the target module by module. For each of the target's FFN layers, SMM aligns the source's grouped FFN layers to it via orthogonal Procrustes, then selects the most informative one to merge in an optimization framework based on the task vector. For the target's attention layers, SMM retains them because of their contextual dependency, a choice supported by the same framework. We evaluate SMM on merging heterogeneous Llama and Qwen models of different sizes. Experiments across language and mathematical reasoning tasks show that SMM helps target models approach source performance while keeping their small model sizes, demonstrating that independently trained heterogeneous language models can serve as reusable carriers for training-free knowledge transfer and enhancement.Our code is available at https://anonymous.4open.science/r/SMM-code-97BC

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.