acceptodds
Under review as a conference paper at ICLR 2027

One Checkpoint, Many Sizes: Quality and Latency in Multi-Exit LMs

Abstract

Multi-exit language models turn one training run into a configurable family of deployment sizes. We characterize their quality and latency through capacity-controlled comparisons of chained Matryoshka suites with independently trained, token-matched model ladders. At the 200M scale, resizing our initially under-allocated later exits (16–31% below their baselines) closes the observed accuracy gap, with panel-dependent gains. At 3B, we evaluate a branched design that combines earlier junction taps with exit-private neck layers. A version within 0.5% of classic largest-exit capacity gains 0.37–0.84 accuracy points over classic in two paired seeds, with lower byte-level perplexity. In these runs, its seven-task accuracy approaches the independent vanilla endpoints and its byte-PPL improves over both. Timing makes the execution tradeoff explicit: deep exits take 22–47% longer per generated token at batch 1 on one H100 despite similar matmul FLOPs. Together, the results show how capacity allocation and branching can improve suite quality, while measured latency identifies the workload-specific cost of each exit. This integrated view of capacity, quality, and runtime provides a concrete basis for designing and selecting multi-exit model families.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.