One Checkpoint, Many Sizes: Quality and Latency in Multi-Exit LMs
Abstract
Multi-exit language models turn one training run into a configurable family of deployment sizes. We characterize their quality and latency through capacity-controlled comparisons of chained Matryoshka suites with independently trained, token-matched model ladders. At the 200M scale, resizing our initially under-allocated later exits (16–31% below their baselines) closes the observed accuracy gap, with panel-dependent gains. At 3B, we evaluate a branched design that combines earlier junction taps with exit-private neck layers. A version within 0.5% of classic largest-exit capacity gains 0.37–0.84 accuracy points over classic in two paired seeds, with lower byte-level perplexity. In these runs, its seven-task accuracy approaches the independent vanilla endpoints and its byte-PPL improves over both. Timing makes the execution tradeoff explicit: deep exits take 22–47% longer per generated token at batch 1 on one H100 despite similar matmul FLOPs. Together, the results show how capacity allocation and branching can improve suite quality, while measured latency identifies the workload-specific cost of each exit. This integrated view of capacity, quality, and runtime provides a concrete basis for designing and selecting multi-exit model families.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.