acceptodds
Under review as a conference paper at ICLR 2027

Selection–Development Lock-in: Access Starvation in On-Policy Learning

Abstract

On-policy learning changes not only the current policy but also the distribution of data available for future learning, a feedback particularly relevant to modern generative-model fine-tuning. A behavioral region may initially perform poorly yet improve with continued exploration, while policy optimization can reduce its future sampling opportunities to negligible levels before its competence develops sufficiently. We call this selection–development lock-in (SDL). In a canonical environment, we formalize and analyze SDL and show that conditional development requires only polynomial effective learning exposure, yet acquiring that exposure through on-policy sampling takes exponential time as initial competence decreases. We extend the analysis to general on-policy dynamics through two quantities: the support lost per unit of conditional improvement and the dependence of development on sampling access. Under explicit regularity conditions, sufficiently steep support loss combined with access-dependent development produces an exponential barrier. We then conduct controlled shared-parameter neural experiments that reproduce the predicted support valley and delayed development beyond the exactly solvable model. These results identify how locally reward-improving selection can make attainable future improvement prohibitively costly within a training budget.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.