Biased Stochastic Gradient for the Log-Softmax Loss
Abstract
The log-Softmax loss is widely used in machine learning, but computing its gradient can be expensive when the number of classes is large. Class sampling reduces this cost, but also raises the concern of bias: whether, after class sampling, the expectation of the stochastic gradient equals the full gradient. Such stochastic gradients have generally been suspected to be biased, but a formal proof has long been missing. Recently, Lin et al. (2025) made an important step toward resolving this question by proving that no estimator based on only a subset of class logits can estimate the full Softmax probability without bias. However, their result concerns Softmax estimation at the logit level only, rather than the gradients with respect to model parameters. Moreover, they only consider class sampling for a fixed training instance, whereas stochastic training involves both instance sampling and class sampling. These differences show that the unbiasedness question itself must be stated more carefully, including which stochastic gradient is considered and which sources of randomness are averaged over. Our first contribution is a proper definition of the unbiasedness of the stochastic gradient for the log-Softmax loss by taking the model, data, parameter values, sampling procedures, and the randomness over expectation into account. With the proposed definition, our second contribution is to formally prove that, for the log-Softmax loss, no universal estimator can ensure unbiased stochastic gradients under class sampling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.