acceptodds
Under review as a conference paper at ICLR 2027

A RIGOR-MATCHED AUDIT OF PERIODIC-STEP LAYER SKIPPING FOR EFFICIENT LLM INFERENCE: CONFLAYERS VERSUS SWIFT, WITH A SUPPLEMENTAL ANALYSIS OF TRAINED ROUTING ALTERNATIVES

Abstract

Layer-skipping methods for efficient LLM inference decide, at some granularity, which transformer layers to execute for a given input. We present a rigor-matched, three-seed audit of two periodic-step, search-based methods that make this decision online, at inference time, and re-evaluate it every few generation steps: a confidence-gated early-exit baseline (ConfLayers) and genuine self-speculative decoding (SWIFT, Xia et al. 2024), together with vanilla autoregressive decoding, across two model scales (Qwen2.5-0.5B and Qwen2.5-1.5B, Yang et al. 2024) and two tasks (GSM8K reasoning, Cobbe et al. 2021; CNN/DailyMail summarization, Nallapati et al. 2016; See et al. 2017). SWIFT is the strongest method on accuracy in three of four cells; ConfLayers is dominated everywhere, with particularly large deficits on GSM8K at 1.5B. Once online-search overhead is correctly separated from pure inference cost — a decomposition we introduce and validate — SWIFT's true inference speed is faster than ConfLayers's in all four cells (5–21%), reversing the naive wall-clock ranking in three of them; ConfLayers's search overhead is small and stable (1–2% of cost) while SWIFT's is larger and considerably more variable seed-to-seed (up to 28.7%). We additionally examine two trained-routing methods, LayerRoute (Sikdar, 2026) (a per-sequence, input-conditioned hard gate) and LayerDrop (Fan et al., 2020) (a fixed, input-independent pruning pattern), as a supplemental analysis rather than a head-to-head comparison, since both operate at a fundamentally coarser decision granularity than the periodic-step methods above. Measuring both under a verified protocol — genuine per-input gating, a genuine full-model baseline, and genuine inference-time compute skipping — both trained-routing methods show modest, real speedups (1.08–1.33x) but accuracy well below the periodic-step methods, including a near-total collapse for LayerRoute on GSM8K at 1.5B (0.003 mean exact-match across three seeds). We release the full audit protocol as a template for rigor-matched efficiency comparisons.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.