Pretrained Aware Merger Lifting: Bringing Pretrained Responses into Model Merging
Abstract
Model merging combines models fine-tuned for different tasks without joint retraining. Task vector values alone do not describe how the shared pretrained model responds to parameter updates. A controlled experiment shows that identical perturbation values can produce substantially different output responses at different parameter locations, revealing information that update values alone do not capture. We introduce Pretrained Aware merger Lifting (PAL), which estimates pretrained sensitivity from a small set of generic unlabeled inputs and constructs an invertible parameter transform. PAL applies the original merging rules in transformed coordinates and maps the merged update back to the original parameter space. The sensitivity estimate and transform require no downstream task data and are computed once per checkpoint for reuse across task sets and mergers. PAL substantially improves recent state of the art mergers across CLIP vision tasks, RoBERTa on GLUE, and Llama-3.2-3B capability merging, with higher aggregate scores for every evaluated base merger.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.