acceptodds
Under review as a conference paper at ICLR 2027

Knowledge That Is Linear Before Fine-Tuning Generalizes After It

Abstract

Not all facts a language model can recall support further learning equally well. We distinguish active facts, which support generalization beyond their own recall, from inert facts, which can be recalled but do not support such generalization in a given task. We study this distinction by constructing new relations from existing ones. Take the relation that maps a place to its continent (Paris → Europe) and assign each of its answers an arbitrary new label (Europe → crab, Asia → zebra). The new relation then maps Paris → crab. We fine-tune the model on the new relation for a subset of subjects, without telling it the original relation or the relabeling rule, and test whether it assigns the corresponding labels to held-out subjects that share an answer in the original relation. Having learned Paris → crab, does it answer Munich → crab? Across eleven language models and 70 relations, how linearly readable the original relation is in pretrained representations strongly predicts this generalization. Within each model, the two correlate across relations at Pearson r = 0.87 to 0.92. This association persists when controlling, separately and jointly, for the number of fine-tuning subjects and answers, the model's accuracy on the original relation, and how often the subjects occur in pretraining text. It is also robust to the arbitrary choices in the construction. Re-randomizing the label assignment and fine-tuning seed, or replacing the animal labels with labels from thirteen other semantic categories, leaves it essentially unchanged. When a fine-tuned model errs, as in Munich → zebra, “zebra” stands for Asia, and Asia is more often among the linear probe's top answers for Munich than among the pretrained model's own top answers to the original question. These findings tie the active/inert distinction to linear readability. Whether a new relation generalizes depends less on whether its facts can be recalled than on how readily their shared structure can be read out linearly.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.