What Survives a Merge? A Controlled Study of Knowledge Allocation and Retention
Abstract
When two fine-tunes of one base model are merged by weight averaging, which of the parents' facts survive? We build a testbed in which provenance is known by construction: two branches fork from one 30M-parameter base, and each memorizes a disjoint set of facts. Under training-template probes, equal-weight merging retains only 29–47% of the exclusive facts that survive at their own endpoints, while the loss barriers practitioners monitor stay near zero. At a fixed total number of exposures per fact, splitting the exposures across branches raises merge survival from 0.23–0.30 to 0.67–0.98 — a training-design lever we call the allocation law. The effect holds on OLMo-2-1B with injected facts (0.34–0.39 to 0.64–0.68, positive on all five seeds) and under open-ended generation. Most of this gain is predicted by averaging the parents' outputs: a null that averages per-fact margins matches the main gap, and one that averages logits also matches survival below threshold; weight averaging loses a further 6–14 points relative to logit averaging. The pattern replicates in fresh synthetic universes, up to 274M parameters, and in true merges of five public fine-tune pairs, where aggregate scores can rise while parent-exclusive knowledge is lost. We release the testbed, per-fact margin fields, and a per-fact ΔNLL screen for deep damage.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.