Grafts Add Up: Scoring Every Layer Swap in One Forward–Backward Pass
Abstract
Layer swapping combines two models fine-tuned from the same base by replacing some layers of one model with layers of the other. Choosing which layers to swap usually requires building and testing many merged models. We study whether the effect of each donor layer can instead be scored on the host model's own trajectory. The scores add across swapped layers; the error is second order when both experts stay near their shared base. SUTURE uses these scores to rank layer windows from one forward and backward pass per probe, then uses a fixed number of built models to calibrate a language constraint and check its choice. Controlled stacks support the approximation for small expert changes and show where its ranking weakens. On public multilingual models, selected windows improve a teacher-forced answer measure over the host but do not surpass published windows. Their effect on the language of free generation remains unmeasured. The method offers a way to narrow layer-swap search while keeping the final choice subject to direct validation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.