acceptodds
Under review as a conference paper at ICLR 2027

Finite-Corpus Dynamics of Supervised and Verifier-Guided Softmax Attention

Abstract

Large-model training commonly separates pretraining and supervised learning from subsequent reinforcement learning with verifiable rewards, yet the dynamical advantage of such stage separation is not well understood theoretically. We study a solvable finite-corpus attention model that combines mean-squared-error (MSE) training with Group Relative Policy Optimization (GRPO). In a single-location regression task, MSE jointly learns a key and value representation, while GRPO uses exact position feedback to refine only the key. In the high-dimensional limit, we derive a dynamical mean-field theory that reduces the finite-corpus dynamics to a self-consistent single-channel process. The theory accurately tracks learning across an abrupt MSE-to-GRPO stage switch and explains verifier-guided refinement after supervised learning. We further compare hard staging with the random interleaving strategy under matched supervision and verifier budgets. Paired experiments reveal systematic schedule crossovers with label noise and sequence length. An exact endpoint decomposition and local force analysis show that these crossovers reflect a competition between clean co-adaptation under interleaving and its continued exposure to noisy supervised updates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.