Where Does Layer-Skip Fragility Live? A Matched-Harness Decomposition of Training-Free Adaptive Layer Skipping
Abstract
Training-free adaptive layer skipping — executing only a subset of transformer layers per token, chosen at decode time by a cheap online signal — is an attractive route to cheaper inference on memory-bound accelerators. In a matched harness on Llama-3.2-1B/3B we ask where skip cost lives and whether a training-free signal can harvest it. Findings are negative for the method, informative for the mechanism. Trailing top-layer early exit is catastrophic (NLL 2.15→5.46 at a 14% layer skip on 3B); interior block skipping at the same budget is several times cheaper. The per-token oracle — skipping exactly the cheapest 25% of positions by a block-skip cost probe — does not upper-bound causal policies (on 3B it is slightly worse than rate-matched random skipping), because per-layer skip costs are sub-additive: summing per-layer costs mis-ranks positions for block skips. Output confidence, the most commonly proposed training-free gate signal, correlates with true skip cost at only r=0.07–0.14 and buys no reproducible advantage over rate-matched random skipping: on 3B it is significantly better than random on one evaluation set (−0.015 nats) and worse on another (+0.003 nats) — a sign flip. A faithful re-implementation of SkipDecode’s actual structure — bypassing lower layers, with a monotonically decreasing entry depth — is far worse than trailing early exit at matched budget (+6.0 vs. +3.3 nats over dense at K=2 on 1B): the depth axis is even more asymmetric than the skip-position axis. Because KV caches are per-layer, a layer skipped at every step can be removed from the decode loop entirely: we measure batch-1 speedups of up to 3.5× (1B) and 1.7× (3B) from true layer omission on A100-40GB — uniform dropping, the least adaptive policy, is the one that converts to wall-clock time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.