acceptodds
Under review as a conference paper at ICLR 2027

Head-Tail Sampled Softmax for Scaling Vocabularies in Language Models

Abstract

Training a language model with softmax cross-entropy scores every vocabulary token at every position, though the loss needs them only for one normalizing sum. On a compact backbone with a large vocabulary, the output layer takes close to a third of a step's FLOPs, 28.9% in Qwen3.5-0.8B. Sampled softmax estimates that sum from a sample, but matches cross-entropy only if each draw is weighted by the inverse of how often it is drawn. Earlier language models trained with in-batch negatives are a sampled softmax without that weight, their scores deflated by token frequency, largely explaining why their direct readout reached 5× cross-entropy's perplexity. We propose head-tail sampled softmax (HTSS), which sums the K most frequent tokens exactly, where sampling is noisiest, and weights the sampled tail. The estimate of that sum stays unbiased at every K, so the loss and gradient consistently estimate those of cross-entropy. HTSS changes only the training loss, so the model and its inference are unchanged, while a step scores a third of the vocabulary. On GPT-2 (124M), weighting alone leaves a 2.1% gap to cross-entropy at matched steps, and at the same cost the exact head cuts it to 1.3%. At equal training compute, HTSS reaches 2.2–2.9% lower perplexity than cross-entropy, and overtakes it on a 304M multilingual backbone from V = 128K. At 256K it reaches 4.5% lower perplexity on 46% less peak training memory, never forming the logit matrix. Finetuning base Qwen3.5-0.8B and Llama 3.2 1B, HTSS lands within 0.34% of cross-entropy's perplexity on 12–20% fewer training FLOPs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.