acceptodds
Under review as a conference paper at ICLR 2027

The Cross-Entropy Gap: Scaling Laws for Sign and Gradient Descent

Abstract

Adam is currently one of the most used optimizers in deep learning, yet our understanding of it remains incomplete. We investigate a simple setting, but one that is rich enough to explain the key advantage of sign-like algorithms such as Adam. We analyze next-token prediction under a linear bigram model with Zipfian token frequencies for the -th token, and study gradient descent and sign descent, a common proxy for Adam, under the cross-entropy loss. Although widely used in practice, the cross-entropy loss remains hard to analyze in theory and exact asymptotic results are rare outside restrictive settings. To address this, we introduce a technique of independent interest: a two-stage reduction that matches discrete gradient descent to gradient flow despite a diverging step-size and replaces the coupled softmax flow by a decoupled exponential flow with closed-form trajectories. For sign descent, we use a fixed step-size that balances errors due to progress and oscillations and model the oscillations as independent uniform variables, which matches experimental results. With this method, we are able to show that gradient descent takes on the order of iterations to reach an error , while sign descent needs only , with only logarithmic dependence on the vocabulary size . This gives a strong advantage to sign descent on language tasks and confirms empirically observed behavior. Our predictions are verified on synthetic data and on OpenWebText.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.