acceptodds
Under review as a conference paper at ICLR 2027

Weights Have Traits: Estimating Fine-Tuning Lineage in LLMs

Abstract

Over a million deep learning models are hosted online, and most are fine-tuned descendants of a few ancestors. Their lineage is rarely fully documented, even though provenance, auditing, and safe deployment depend on it. We hypothesize that fine-tuning acts as an evolutionary process, so that weights have traits inherited from their ancestors. We show that a model phylogeny (evolutionary tree) can be estimated from the weights of the observed leaf models alone, with all ancestors latent. On controlled LLM fine-tuning trees whose true structure is known, our estimates contain every true clade (a group of models descended from one ancestor) in 85–100% of trees, compared with 3% for a random tree. Treating each weight layer as a gene-like estimator of lineage, we find that a single weight matrix, such as one attention key matrix (0.1–0.2% of the weights), estimates the phylogeny as well as all the weights combined. Moreover, estimation accuracy is positively associated with two measures from phylogenetics, the Atteson margin and a tree-likeness score based on four-point additivity. The latter needs no true tree. We also report two secondary findings. Models closer in weight space often behave more similarly on held-out tasks, and weight-based estimation matches behavior-based estimation (PhyloLM) on two documented real families while outperforming it on the controlled Llama trees (100% of clades versus 35–56%).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.