acceptodds
Under review as a conference paper at ICLR 2027

Token Averaging for Compute-Efficient Language Model Pretraining

Abstract

Training language models on more text ordinarily requires proportionally more compute. We introduce token averaging, a parameter-free method that replaces each group of consecutive token embeddings with their mean before the first attention layer. The transformer therefore processes times fewer positions while predicting the first token of the next group. This allows a fixed training budget to cover approximately times more raw text; at , the model trains on twice as much text using - fewer FLOPs than the baseline (without token averaging). Since token embeddings are averaged, we randomize group boundaries during training and evaluate using different boundary offsets, ensuring that the token-averaged model and the baseline are scored on identical target positions. In each offset, a sequence times shorter is processed, so the cost of all offset evaluations is approximately equal to one forward pass for the baseline. At inference, each pass stores as many KV-cache entries and incurs of the attention cost compared with the baseline, i.e., half the cache and one-quarter of the attention cost at . Across models from 50M to 500M parameters, our experiments show that the baseline requires , , , and more compute, respectively, to reach the same loss as the model. Full-position evaluation shows lower perplexity at 125M with fewer training FLOPs and comparable loss at 500M with approximately fewer FLOPs. Additional experiments show that learned pooling over token content and within-group position outperforms mean pooling by nats at and nats at . Together, these results show that token averaging increases both training-data throughput and inference efficiency without compromising model quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.