acceptodds
Under review as a conference paper at ICLR 2027

Same Knowledge, Different Use: How Training Pairings Shape Compositional Generalization in Language Models

Abstract

Fine-tuning can teach a language model to recall a fact without teaching it to apply that fact in a new problem. We study how the assignment of training examples contributes to this gap. Models first learn a recorded symbol for each of 384 fictional entities and learn to classify six-symbol sequences when all inputs are supplied. An application question supplies only the first five symbols and names the entity whose record specifies the sixth. The target is a category label denoting a range of the final-state probability under a fixed two-state hidden Markov model. We compare two training conditions using identical starting weights within each comparison. Each of 96 application-trained names receives eight different five-symbol contexts. In one condition, all eight contexts require the same answer for that name: SAME8 (8:0). In the other, four require one answer and four another: CROSS8 (4:4). We construct this difference by exchanging contexts between names with the same recorded symbol. The overall application material, labels, number of contexts per name, factual and explicit-task replay, and training updates remain matched. On contexts withheld from the compared application training, 4:4 improves accuracy by an average of 36.8 percentage points across three Qwen worlds, 38.4 points in Llama, and 39.8 points across two Phi assignments sharing one starting checkpoint and world. These gains do not accompany better factual recall. A subsequent Qwen comparison shows that 7:1 training already produces 29–36-point gains over 8:0. Its single minority-answer context is presented six times, and equivalence to 4:4 is not established. The original 4:4 advantage persists when both conditions' application labels uniquely identify the recorded symbol. Transfer to names without application examples is less consistent, and better ordinary application can coexist with worse performance when instructions require using a conflicting supplied symbol. Replacing selected feed-forward components between the original trained copies partially improves answers without identifying a unique composition algorithm. The results show that how the same available application material is assigned can substantially change generalization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.