Dynamic Depth Routing of Large Language Models with Mixture of Segment Operators
Abstract
Large language models (LLMs) process every input through a deep stack of layers, which makes inference slow and costly, yet many inputs do not need the full stack. Existing methods reduce this depth either by permanently pruning layers or by learning to skip layers per input, but most decide about one layer at a time, even though the quality of the reduced model depends on the whole schedule of which layers are skipped, run or reused. We propose MOSO (Mixture of Segment Operators), which casts depth reduction as learning a routing schedule over a frozen pretrained model. Its actions are segment operators that skip, run or repeat a short group of consecutive layers; this family contains static pruning, layer skipping and looped computation as special cases, and, under explicit linear-block conditions, we show that a segment repeat can realize maps that no layer-by-layer action sequence can. A lightweight router scores all candidate segments jointly, so each decision is made in the context of the schedule being built, and the schedule is learned end to end without any search over schedules. Across five LLMs from three model families, MOSO attains the highest average accuracy at each reported operating point, with its lead over the strongest baseline growing at the tighter one, and controlled experiments show that the learned schedule contributes to it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.