acceptodds
Under review as a conference paper at ICLR 2027

MOMA: Masked Orthogonal Matrix Alignment for Zero-Additional-Parameter Model Merging

Abstract

Model merging offers a scalable alternative to multi-task learning but often suffers from substantial performance drops on classification tasks. We attribute this degradation to a geometric misalignment between the merged encoder and static task-specific classifier heads. Existing post-merging methods rely on auxiliary parameters to enforce strict representation alignment, incurring storage and inference overhead that grows with the number of merged models. We challenge this approach by revealing that the misalignment is well-approximated by an orthogonal transformation, rendering such strict alignment unnecessary. Leveraging this insight, we propose MOMA (Masked Orthogonal Matrix Alignment), a plug-and-play post-merging framework that rectifies the misalignment by jointly optimizing a global multi-task vector mask and task-specific orthogonal transformations. Crucially, MOMA absorbs corresponding new parameters directly into existing model weights, achieving state-of-the-art performance among post-merging methods with zero additional parameters and zero added inference cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.