Merging Destroys What You Paid to Teach
Abstract
Large Language Models (LLMs) are increasingly specialized by fine-tuning copies of a shared base model and then merging them into a single model. While model merging is typically evaluated by how well the merged model preserves downstream benchmark performance, these evaluations need not reveal whether the information acquired during specialization survives the merge. We demonstrate this gap on multilingual and coding specialists: merged models with similar benchmark performance can retain very different amounts of the knowledge acquired during specialization. To study this failure systematically, we introduce a suite of ten factual datasets designed to measure knowledge acquired during fine-tuning. Across these datasets, factual retention drops substantially after merging, worsens as more specialists are combined, and can be highly uneven across specialists. To better preserve this knowledge, we introduce LUMEN (Low-rank Update Merging with residual ENhancement), which jointly learns a coordinate-wise mixture of specialist updates and a low-rank correction for the merged model using only a small subset (%) of specialist data. Across merges of two to six specialists, LUMEN improves normalized factual retention by % on average over the strongest reported baseline at each merge size. At six specialists, LUMEN retains % pre-merge factual accuracy, compared with % for the strongest reported baseline. Together, our results establish factual retention as a distinct objective for model merging and show that substantially more of the knowledge acquired during specialization can be preserved at no additional inference cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.