acceptodds
Under review as a conference paper at ICLR 2027

The Delta Is Not the Skill: Zero-Prior Elicitation Transfers LoRA Behavior Across Model Families

Abstract

Reusing low-rank adapters (LoRA) across model sizes, generations, and architectural families is increasingly essential for deploying specialized capabilities. However, native retraining requires inaccessible training data, while existing synthetic-data transfers rely on task-specific priors (task names, seed examples, labels, or auxiliary discriminators). Weight-space transfer avoids these priors, but we uncover a sharp dichotomy: within a shared parameter basis, zero-edit adapter transplantation is near-lossless (recovering – of native gains), whereas across distinct families, weight-space routes collapse () because low-rank deltas are base-relative. What survives across model boundaries is behavior, not weights. We introduce *zero-prior skill elicitation*, a framework that extracts adapter capabilities without task identities, seed data, labels, discriminators, or white-box access. By identifying inputs where adapter and base generations diverge, evolving these probes to broaden coverage, and distilling the adapter-equipped teacher's completions into students, we transfer complex skills purely through text. Distilling from a single 397B MoE teacher LoRA matches or exceeds native supervised fine-tuning across four student models in two architectural families, reducing target training compute by –. Gains are pronounced on multi-step reasoning (elevating 9B MATH-500 by points; ), competition mathematics, and out-of-distribution instruction following, with zero verbatim -gram benchmark leakage and successful generalization to Granite models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.