acceptodds
Under review as a conference paper at ICLR 2027

Unified Merging of Heterogeneous Pretrained Models: From Layer Selection to Semantic Alignment

Abstract

As a novel technique for large pretrained models, model merging integrates multiple independently pretrained or fine-tuned models to construct a unified and efficient model without accessing the original training data or incurring an expensive computational cost. Existing model merging methods predominantly focus on parameter-level fusion within a fixed model architecture, facing inherent limitations in scalability and performance enhancement. While merging models at the architecture level shows promising potential, current approaches lack a clear and unified problem formulation, and suffer from exponential growth of the search space, and incompatibility of inter-layer feature representations. To address these challenges, we propose a unified architecture-parameter merging framework named UniMerge. Specifically, UniMerge formulates model merging as a Markov decision process and incorporates parameter-level merged candidates into an architecture-level search pool. To solve the formulated problem, UniMerge designs a state-enhanced embedding network to capture inter-layer parameter correlations and introduces a contrastive self-supervised learning-assisted actor-critic policy to improve exploration efficiency. To mitigate inter-layer incompatibility, we develop a glue layer based on dual knowledge distillation to ensure inter-layer semantic alignment and cross-layer feature representation. Extensive experiments on language and vision models demonstrate strong performance and search efficiency, while cross-scale experiments further show that UniMerge can stitch compatible models with different widths and depths, extending architecture-level merging from fine-tuning-induced heterogeneity to explicit structural heterogeneity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.