When Does Orthogonal Training Help Model Merging? Operator-Relative Disentanglement in Task Arithmetic
Abstract
Task arithmetic combines independently fine-tuned models through their parameter updates. Training methods and merge rules are usually assessed one component at a time, leaving open whether a training intervention benefits merge rules uniformly or changes their relative utility. We study orthogonality-regularized training under a stricter test: after allowing every merge rule to improve or deteriorate together, does training still change their gains over direct addition? We formalize this question as operator-relative disentanglement and test multiplicity-controlled differences between operator responses. Across image and language task suites, orthogonal training changes relative operator responses rather than producing only a common gain shift. Under our prespecified inference procedure, two vision backbones exhibit resolved preference reversals, a language suite exhibits resolved changes in relative gaps, and a larger vision backbone provides a null result under the primary metric. Reaggregating the same frozen outputs using worst-quartile rather than average utility also changes the response pattern. Shared-radius searches, fixed-map equal-norm audits, validation-only selection, and identity-inclusive deployment separate compatibility from global scale and endpoint quality. A local second-order analysis explains why Euclidean orthogonality alone cannot guarantee compatibility after a merge rule transforms the updates. These results motivate evaluating training, merging, and functional metrics as a coupled system.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.