The Retention Gap: Why Optimal Features Need Not Remain Stationary
Abstract
Implicit-bias analyses often explain selection by asking which features training captures. A separate question is whether captured optimal features remain dynamically stable as the residual evolves. This paper isolates that distinction in an explicit two-dimensional regression problem whose minimum representation norm and complete optimal contact geometry are known exactly. An occupied optimal dual contact can lose stationarity as training proceeds; this dynamical failure is termed the retention gap. For a compact family of states already resting on the optimal contacts, a computer-assisted proof certifies a uniform transversal first exit with positive charge, and an exact local calculation shows immediate motion into positive dual slack. An intrinsic canonical-measure identity then connects this local event to function selection: at interpolation, excess representation norm equals canonical mass weighted by signed dual slack. Local loss of retention and persistence to the endpoint are therefore distinct obligations. Complementary full-width experiments lie outside the certified entrance family and do not test its event ordering. Instead, inspected trajectories exhibit the same qualitative sequence of approach, margin loss, and departure at the other contact. Across the robustness sweep, training reaches near-machine-precision loss with a positive endpoint excess diagnostic and a consistent off-contact signature, while splitting the exact contact tie changes the observed scale dependence. Overall, these results show that reaching an optimal feature geometry does not by itself make that geometry dynamically stable, identifying retention as a distinct requirement in feature-capture accounts of implicit bias.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.