acceptodds
Under review as a conference paper at ICLR 2027

Gated Task Arithmetic: Controlling Task-Specific Input Information Flow in Model Merging

Abstract

Model merging has emerged as an effective approach for consolidating multiple task-specific fine-tuned models into a single unified model. Among existing approaches, data-free merging is particularly attractive as it does not require access to the original training data. A central challenge in model merging is to preserve task-specific knowledge while mitigating interference among task vectors. However, existing merging methods mainly manipulate task vectors themselves, without explicitly controlling the input representations on which they operate. After merging, each task vector may therefore operate on input representations associated with other tasks, potentially introducing cross-task interference. In this work, we propose **Gated Task Arithmetic (GTA)**, a data-free model merging method that introduces a task-specific input gate for each task vector. GTA constructs a gated task vector as and performs task arithmetic over the resulting gated task vectors. We design as a **task-preserving and interference-suppressing gate (PSI-Gate)**, which preserves input representations associated with its corresponding task while suppressing those associated with other tasks. To learn the PSI-Gates without the original training data, we introduce a **task-preserving and interference-suppressing surrogate loss (PSI surrogate loss)** using task-vector-derived surrogate input representations. Under task-vector Gram covariance estimates and the stated convergence conditions, finite-step gate optimization from identity induces a preservation-oriented spectral bias, while its limiting solution corresponds to closed-form activation-aware merging. We extensively evaluate GTA across ViT models with varying numbers of vision tasks, as well as LLaMA and Qwen2 models fine-tuned on diverse natural language tasks. GTA consistently achieves strong merging performance across different model architectures, scales, and modalities, improving the recovery rate of task-vector performance by more than **16 percentage points** over existing merging methods in the best case.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.