acceptodds
Under review as a conference paper at ICLR 2027

A Data-Centric Approach to Improving LLM Generalization

Abstract

Simplicity bias (SB)—the tendency of gradient-based methods to prioritize dominant features—shapes how deep networks learn and generalize. Despite its demonstrated impact on generalization in CNNs, the role of SB in transformer generalization remains largely unexplored. We theoretically analyze a multi-head self-attention model and prove that Sharpness-Aware Minimization (SAM) reduces simplicity bias in transformers. Together with the empirical observation that SAM improves generalization, this suggests that reducing SB may contribute to improved generalization. Our theory further reveals that examples associated with dominant features exhibit rapid early loss drops, while examples associated with less-represented features lag behind. This makes SB observable through training dynamics and motivates a simple data-centric approach: we partition data based on early loss dynamics and amplify less-represented examples through upsampling or synthetic generation. For synthetic generation, we find that generated examples should remain sufficiently similar to their corresponding real data, achievable with diffusion LLMs. Across Llama2 (7B), Phi2 (2.7B), Llama3.2 (1B), Gemma3 (1B), and Qwen3 (0.6B), our approach consistently improves performance with AdamW and Muon, achieving a 6x wall-clock speedup and 2x memory reduction over SAM. Improvements reach 8.3% on mathematical reasoning, 26% on code generation, and 4.1% on instruction following.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.