Adaptive DCT-Domain Spectral Merging: Rethinking Frequency-Domain Model Merging
Abstract
Model merging combines task-specific fine-tunes of a shared base into one model without retraining or training data. Frequency-domain methods such as FR-Merging reduce cross-task interference by applying a fixed-cutoff 2D FFT high-pass filter to each task vector before combining them. We revisit this recipe along two axes: replacing the FFT with the discrete cosine transform (DCT), which better matches the real-valued, non-periodic structure of weight matrices, and replacing the global cutoff with a per-layer adaptive cutoff driven by a cheap, data-free cross-task sign-interference score. On MergeBench, however, we find that two combination-rule pitfalls dominate the transform choice: an uncalibrated edit-scale that produces larger updates than the baselines at matched layers, and energy-proportional weighting under which each task's contribution scales with the square of its task-vector norm, starving smaller-update domains (instruction-following receives under 6% of the combination weight versus an equal 20% share). Correcting both raises aggregate normalized performance from 0.806 to 0.966 on Llama-3.1-8B-Instruct, matching DARE (0.965) on the four domains we aggregate. A replication across two model families at four scales (Llama-3.1-8B, Llama-3.2-3B, Gemma-2-2B/9B; 36 spectral merges, 144 domain evaluations) confirms the combination rules as the dominant effect: edit-scale calibration lifts pooled four-domain scores by +0.11 on Llama-3.1-8B and +0.07 on Gemma-2-9B and is neutral to slightly negative on the two smaller bases, versus at most 0.015 for any DCT-vs-FFT contrast; the DCT is ahead in most pairwise comparisons, but by margins we do not claim as significant. We release the corrected method, the diagnostics, and comparisons against Task Arithmetic, TIES, DARE, and an FR-Merging-style FFT baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.