Can LLMs Learn to Generalize before Memorizing High-Order Markov Chains?
Abstract
Early stopping is a widely-used implicit regularization technique that halts training before full optimization of training to avoid overfitting. In supervised learning it is an effective safeguard against overfitting, and it has recently been identified as a key ingredient of generalization in diffusion models. We ask whether early stopping can similarly prevent memorization when causal language models are trained on sequences generated by order- Markov chains. Our answer is mixed. For low-order chains, early stopping leads to generalization, but for high-order chains it is not sufficient. Theoretically, we establish a regularization sufficient for learning order- chains, which implicitly holds early in training. Empirically, we find that for high-order chains this implicit regularization vanishes early in training, so stopping fails to prevent memorization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.