On the Edge of Understanding: Churn in LLM Post-Training
Abstract
Beneath steady accuracy improvements in LLM post-training, prior work finds that many queries flip between solved and unsolved across training epochs. However, the underlying mechanisms driving these flips remain poorly understood. Through large scale analysis, we find a single gradient update flips of learnable queries between solved and unsolved throughout training, across both in-distribution and out-of-distribution datasets — a phenomenon we term . Accuracy gain is the surplus of positive over negative churn, both persisting even after accuracy plateaus. Churning queries reside on LLM's edge of understanding, razor-thin confidence margin between correct and incorrect answers makes correct answer more vulnerable to perturbation. Persistent churn incurs substantial cost: solved queries flips to unsolved lead to pp forgetting gap, and of final accuracy comes from unreliable answer on queries that flips repeatedly across training checkpoints. Finally, we find that churn is not benign: applying sharpness-aware minimization after accuracy plateaus reduces churn while lifting final accuracy by percentage points for models up to parameters, with no benefits on larger models. Our work highlights the importance of analyzing the evolution of query-level performance during post-training and points to a promising path for improving learning efficiency via reduced churn.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.