acceptodds
Under review as a conference paper at ICLR 2027

Merge by Teaching: Consolidating a LoRA Library by Distillation

Abstract

Fine-tuned adapters accumulate. A team that has collected a dozen LoRA modules, one per behavior, trained by different people at different times, eventually wants them in a single model, and the field’s default answer is to add their weight deltas together. We show that this answer quietly stops working. As a library grows, summed adapters do not merely blur. At thirty-two skills the merged model scores zero on every one of them, and on a knowledge benchmark none of the adapters ever touched it answers no better than a coin, because the deltas disagree in sign about as often as coin flips. Yet nothing is actually lost. Recent post-training work has established the remedy in a different setting — let each specialist teach the student on the student’s own outputs rather than donate its weights — and we ask what happens when that operator is taken out of the lab that trained the specialists and into a library someone inherited. Three things change, and each is a result. Nobody knows which adapter owns which prompt, so the router has to be learned. The library is far larger than the language-model studies of this operator, which is exactly the range where weight arithmetic dies. And the base model’s own abilities, which that work protects but never measures, turn out to be collateral damage for weight arithmetic and imitation alike. Teaching leaves them standing, and its one loss has a nameable cause. Where the sum scores zero on all thirty-two skills, the student keeps 92% of what the teachers reach alone. We call the recipe Merge by Teaching, and we report where it holds, over three model families, five sizes, and adapters we did not train, along with the one failure mode we can characterize but have not yet learned to prevent.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.