acceptodds
Under review as a conference paper at ICLR 2027

The Durable Asset Is the Data, Not the Weights

Abstract

Organizations increasingly serve fleets of fine-tuned specialists on open-weight language models whose bases improve every few months. Each release reopens the same decision: freeze the old specialist, port its weights, refresh it from retained behavior, or retrain. These strategies form a ladder of retained resources—nothing, weights, inputs, labels—yet the decision has never been measured: prior work evaluates ad-hoc model pairs, never a real release sequence. We measure it longitudinally, on a complete fourteen-month release lineage (four consecutive generations, two scales, six tasks spanning classification, structured generation, and tool use), validated on a second family with fully documented training genealogy; hypotheses are fixed before each experimental wave, and every comparison is paired, per-example, and cost-accounted. We find that (1) specialization is durable: a frozen specialist can outperform the newest base for over a year, governed by a measurable task–base coupling rather than task class; (2) weights are fragile: porting between architecturally identical checkpoints from independent pretraining runs collapses below the no-adapter floor, yet porting onto a short continued-pretraining descendant is nearly free—portability is a distance budget in pretraining tokens, not an architecture or family property; and (3) data recovers what weights cannot carry: distilling the old specialist into the new base, using retained inputs only, matches gold-label retraining with zero new annotation. A pre-specified policy built on these facts nearly matches always-retraining at a third of its cost with zero regressions, while the shape-compatibility rule practitioners use is the worst policy we test. All code, records, and cost logs are released; everything runs on one consumer GPU.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.