acceptodds
Under review as a conference paper at ICLR 2027

Recall, Reason, Regress: How Reasoning Data Mixtures Shape Generalization in LLMs

Abstract

How does increasing the share of reasoning supervision change an LLM’s ability to use what it knows? We investigate this question in a controlled relational reasoning framework that separates fact recall, within-pattern generalization, and cross-pattern transfer. Across four model sizes (4B–32B) and 15 nonzero data mixtures, we vary the number of reasoning demonstrations under a fixed training-row budget while holding the demonstrated composition patterns constant. We identify three empirical regimes: Recall, Reason, and Regress. Knowledge-only training supports perfect fact and rule recall but fails to produce correct composite answers. Sparse reasoning demonstrations bridge this gap, enabling generalization to new instances and undemonstrated compositions, with within-pattern errors decreasing approximately as a power law. Yet further increasing supervision can reverse cross-pattern gains while within-pattern accuracy remains near saturation: doubling the reasoning-data share from 6.4% to 12.8% reduces cross-pattern accuracy by 8.82–20.31 percentage points across all four models, with within-pattern changes of at most 0.16 points. This decline persists when independently probed prerequisites remain accessible; pattern-specific shifts in generated solutions are consistent with specialization to demonstrated patterns. At a common training endpoint, model scale raises observed peak transfer from 55.94% at 4B to 96.74% at 32B, but the grid-maximizing mixture does not vary monotonically with model size. These findings make cross-pattern transfer an essential objective for selecting reasoning mixtures: continued gains on demonstrated patterns can conceal a narrowing of compositional generalization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.