acceptodds
Under review as a conference paper at ICLR 2027

AVERT: Average-Depth Variable Execution via Request-Local Token Routing

Abstract

Layer skipping can reduce large language model inference costs, but a request budget does not specify how to allocate layer execution across generated tokens. We introduce AVERT, which treats a continuous request budget as a target for mean executed depth. At each dynamic layer, a request-local router uses the current hidden state and that request’s budget to decide whether to execute or skip the Transformer block, allowing token depths to vary without coupling decisions across requests. Skipped tokens bypass the whole Transformer block, omit their layer-local key/value (K/V) entries, and use a lightweight residual-repair adapter. A single checkpoint serves requests with different budgets in the same batch. Across three LLaMA backbones, six quality metrics, and 13 requested budgets, AVERT is best or tied for best in 206 of 234 comparisons. Across homogeneous and mixed-budget batches and five off-grid targets, the mean executed-layer fraction for each target differs from the requested budget by at most 1.15 percentage points. When splitting requests by budget leaves batches underfilled, mixing budgets reduces decode time per output token (TPOT) by 4.3–21.2%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.