Understanding Sparsification in Multi-Task Model Merging via Linearized Functional Decomposition
Abstract
Multi-task model merging offers an efficient alternative to multi-task training, yet post-hoc parameter sparsification remains weakly explained by training-time analyses. We introduce linearized functional decomposition (LFD), which separates target-task and non-target contributions and bounds the merged–LFD discrepancy. The resulting path diagnostics also track LFD–target consistency and retained single-task quality. We examine how transformations change task-vector norms and cross-task interactions, then compare the bound with measured path behavior. Across vision and language tasks, effective sparsification mainly narrows the merged–LFD gap, while the other diagnostics reveal changes in target consistency and capability. The diagnostics also predict relative gains for held-out tasks and show where transformations beyond sparsification help or hurt.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.