Training Dynamics of Large State-Transition Models
Abstract
Jointly learning input and output embeddings is common practice in modern neural sequence models, yet it remains unclear when gradient-based training can successfully optimize these coupled components to achieve accurate predictions. We study this question in a softmax state-transition model with a learned state embedding and prediction head, trained from Gaussian initialization using exact population gradients. We consider a model with states and embedding dimension , where the ground-truth transition law’s row-centered log-probability table has low rank and bounded entries. We measure prediction quality by relative prediction gain (RPG), normalized so that uniform prediction has gain zero and optimal prediction has gain one. We provide analysis for both untied and tied models. With an untied model and weight decay of order , we show that perturbed gradient descent achieves RPG above in polynomial time when . With probability at least where , this holds at every iterate in the final fraction of the run while perturbations remain at a fixed scale. For tied embeddings, we provetraining guarantees on the observed current states, and under the assumption of ground-truth feature coverage and bounded learned rows, we prove prediction guarantees on unseen current states. We present numerical experiments with practical parameter choices that qualitatively illustrate persistent training gain and tied-model transfer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.