TOKEN MARKOV GRAPHS FOR DATA-EFFICIENT LANGUAGE MODEL PRETRAINING
Abstract
The cost of pretraining a language model is set by the tokens, and the wall-clock time, needed to reach a target loss, and high-quality text is becoming the binding resource. We introduce the Token Markov Graph (TMG), a directed graph of token-transition counts computed once from the pretraining corpus, and inject it into pretraining as an annealed graph-distillation loss or as a lightweight architectural memory module. TMG reduces the tokens needed to reach the baseline’s final validation loss at and below the Chinchilla-optimal budget. At the Chinchillaoptimal budget, the architectural route (engram v1) needs 8.1% fewer tokens and 5.1% less wall-clock time, with a gain in all 9 seeds (t(8) = 7.61, 4.56, and 3.05 for tokens, wall-clock time, and final loss gap). At a fixed 100M-token budget, curvature-weighted distillation saves 14.2–20.0% of tokens for models from 172M to 3B parameters (minimum t = 4.83 on the loss gap), with token savings, and wall-clock savings of up to 20.9%, growing with model size. The gains hold under a tuned baseline learning rate, on two independently built graphs, and on code-only training (6 seeds, t(5) = 3.96), and carry over to next-word cloze completion. Because the graph is built once and reused across models and runs, corpus transition statistics offer an inexpensive, broadly applicable accelerant for pretraining.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.