MinimaxMerge: Balancing Task Retention in Training-Free Model Merging
Abstract
Training-free model merging combines task-specialized LLMs without additional data or training, yet existing methods largely focus on reducing interference among task vectors while optimizing aggregate performance, leaving the weakest task unprotected. We introduce MinimaxMerge, a training-free and data-free approach that instead optimizes a worst-task reconstruction objective when combining multiple specialists. Across mathematics, code, and science experts, MinimaxMerge achieves the highest worst-task retention on dense Gemma-4-12B (1.0067 vs. 1.0048 for the strongest baseline) while maintaining mean performance. On Gemma-4-26B-A4B, under attention-only adaptation, it achieves the highest worst-task accuracy, while TIES narrowly outperforms it in worst-task retention (1.0196 vs. 1.0187). Replacing the minimax objective with a mean objective decreases worst-task accuracy by 1.66 points, demonstrating that objective choice materially affects retention beyond the merging procedure itself. Cross-family experiments further show that these gains are not universal: TIES outperforms MinimaxMerge on OLMo-2-13B (0.9672 vs. 0.9542 worst-task retention), while the mean objective performs best on SmolLM3-3B. Together, these results establish the merge objective as an important but model-dependent design dimension for preserving capabilities in multi-task model merging.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.