acceptodds
Under review as a conference paper at ICLR 2027

Matryoshka Language Model Suites

Abstract

Training a language model suite classically requires training each model separately and serving them independently. We improve both training and inference efficiency by stacking sub-models of increasing size into a single nested architecture trained end-to-end. This Matryoshka training framework reduces the total parameter count of the suite, enables low-cost distillation from the largest to all smaller sub-models at every training step, and is well-suited for speculative decoding as the draft model is contained within the verifier. We validate our approach by training a Matryoshka suite comprising 500M, 1.5B, and 3B sub-models. Our suite is on par with independently trained baselines on validation perplexity, out-of-domain perplexity, and benchmark performance while using 36% less training compute, and achieves 14-26% higher throughput via speculative decoding. We also ablate key architectural choices, offering guidance for building strong Matryoshka LM suites.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.