Generalization Before Memorization: How Early Objective Prioritization Shapes Subsequent Learning
Abstract
Large language models (LLMs) both learn abstract patterns that generalize to unseen examples and memorize specific knowledge that cannot be inferred from patterns. Although models typically learn both objectives concurrently, we identify a fundamental learning order asymmetry: prioritizing generalization first leads to stronger subsequent generalization, whereas prioritizing memorization first can make generalization harder to acquire. Across 54 model–dataset configurations spanning six architectures and nine tasks, generalization-first training improves generalization performance in 74.1% of cases while preserving memorization. Conversely, a memorization-first regime frequently prevents models from developing skills beyond memorizing data. To further investigate this finding, we conduct an analysis of intermediate neural representations across training, and isolate specific sub-networks responsible for this behavioral divergence. Our findings demonstrate that training order governs optimization trajectories, highlighting the need to prioritize generalization early in LLM training schedules.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.