Finite-Corpus Dynamics of Supervised and Verifier-Guided Softmax Attention
Abstract
Large-model training commonly separates pretraining and supervised learning from subsequent reinforcement learning with verifiable rewards, yet the dynamical advantage of such stage separation is not well understood theoretically. We study a solvable finite-corpus attention model that combines mean-squared-error (MSE) training with Group Relative Policy Optimization (GRPO). In a single-location regression task, MSE jointly learns a key and value representation, while GRPO uses exact position feedback to refine only the key. In the high-dimensional limit, we derive a dynamical mean-field theory that reduces the finite-corpus dynamics to a self-consistent single-channel process. The theory accurately tracks learning across an abrupt MSE-to-GRPO stage switch and explains verifier-guided refinement after supervised learning. We further compare hard staging with the random interleaving strategy under matched supervision and verifier budgets. Paired experiments reveal systematic schedule crossovers with label noise and sequence length. An exact endpoint decomposition and local force analysis show that these crossovers reflect a competition between clean co-adaptation under interleaving and its continued exposure to noisy supervised updates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.