Next Concept Prediction Leads to Stronger and More Scalable Language Models
Abstract
We propose Next Concept Prediction (NCP), a generative pretraining paradigm built on top of Next Token Prediction (NTP). NCP explicitly predicts higher-level concepts spanning multiple tokens, introducing an additional concept-level pretraining objective and using the predicted concepts to guide subsequent token generation. Our model, ConceptLM, constructs a learned concept vocabulary from its hidden states through vector quantization. A dedicated Concept Module predicts future concepts on the constrained semantic manifold defined by this vocabulary, with NCP and NTP trained jointly end-to-end. We first validate the basic framework on GPT-2 and Pythia backbones, then extend it to trillion-token pretraining on the fully open OLMo-3 architecture. With hierarchical residual routing and an adapted optimization recipe, we train an 8.9B-parameter model from scratch entirely on 5.73T tokens of open Dolma-3 data. The resulting model reaches the final pretraining loss of OLMo-3-7B using only 51.3% of its training tokens and improves the downstream average by 2.45 points, including a 5.99-point gain on GSM8K. The learned concept pathway also supports lightweight domain adaptation. Together, these results demonstrate the potential of joint token and concept prediction for scalable language-model pretraining with open data and model architectures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.