StrataSpec: Self-Speculative Decoding with Structured Routes for Efficient LLM Serving
Abstract
Efficient LLM serving under limited compute depends on batching, which amortizes each model pass across several requests. Self-speculative decoding (SSD) reduces proposal cost further without a separate draft model by drafting through a route, a subset of the target model’s attention and MLP sub-layers, before full-model verification. Existing methods adapt this route to each request from verifier feedback, but assess each choice in isolation. A shared proposal pass executes the union of those routes. Because independently chosen routes rarely skip the same sub-layers, their union can reactivate almost the full model and give different row sets to its sub-layers. Much of the per-request saving then disappears at the batch level. We introduce StrataSpec, which retains request-specific route choice within one compact, regular proposal pass. Its routes form a nested family in which every richer route contains the preceding one. The union of any mixture is exactly its richest assigned route, and sorting requests by route gives each sub-layer a contiguous suffix of rows. A batch planner chooses routes and draft length from the incremental work they add to the active batch. StrataSpec improves throughput by up to 4.40× over the strongest of three recent adaptive SSD baselines and sustains 98.7–99.3% SLO attainment under continuous arrivals, while baseline queues grow by orders of magnitude.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.