acceptodds
Under review as a conference paper at ICLR 2027

When Grokking Fails: Phase Boundaries under Data Skew and Low-Dose Rescue

Abstract

Grokking is usually studied under fixed training distributions, leaving unclear how data imbalance changes whether delayed generalization is reached at all. We study this question by continuously varying data skew and weight decay in an autoregressive counting task. Increasing skew systematically contracts the region in which models grok and reveals two qualitatively different failure regimes: a boundary regime where additional training can recover generalization, and a deep-trap regime that remains ungeneralized even after 10^6 training steps. This separation shows that skew can change the accessibility of generalizing solutions, rather than merely delaying their emergence. We then ask whether this loss of accessibility can be reversed through small changes to the training distribution. Replacing fewer than 1% of minibatch slots with samples drawn from a broader distribution can restore grokking, while sample-count-matched interventions that preserve the original sampling distribution or target a narrow rare-state pool are substantially less effective. Rescue strengthens as the support of the injection distribution broadens and remains possible even when putative bottleneck states are excluded, arguing against a simple missing-rare-example explanation. The same intervention ordering transfers to modular addition. On Dyck-1, skew again produces a structured failure region and broader injection remains preferable, although complete rescue does not transfer and longer training can expose late-stage collapse. Overall, these results identify training-distribution coverage as an empirical control axis for both the accessibility and recovery of grokking under data skew.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.