acceptodds
Under review as a conference paper at ICLR 2027

WHAT MAKES A LAYER IMPORTANT? DISSECTING LAYER–IMPORTANCE SIGNALS AND LEARNING CONDITIONAL LAYER SKIPPING

Abstract

Which layers of a transformer matter, and can cheap signals tell us? We first present a systematic cross‑architecture measurement of three popular layer importance signals—attention entropy, hidden‑state norm, and gradient norm—against exhaustive single‑layer ablation as ground truth, on GPT‑2 (causal) and BERT (bidirectional) across five NLP tasks. The result is a consistent negative finding with identifiable failure modes: attention entropy is highly consistent across tasks (mean pairwise r=0.99 on GPT‑2, 0.95 on BERT) yet uncorrelated with importance (r=0.37 collapsing to -0.15 when a single high‑leverage layer is removed on GPT‑2; r=0.02 on BERT); hidden‑state norm tracks depth only under pre‑LN residual streams () and is rendered uninformative by post‑LN normalization; and gradient norm's apparent success (r=0.95 on BERT) is a single‑outlier artifact (). We formalize these failure modes with three theoretical propositions. Motivated by our diagnosis, we introduce LIFT (Layer‑Importance Fusion for conditional layer skipping), a lightweight gating module with less than 0.3% additional parameters. LIFT fuses multiple weak importance signals conditioned on input representations, trained via straight‑through Gumbel‑Softmax with knowledge distillation and compute‑budget regularization. Experiments demonstrate that LIFT substantially outperforms heuristic baselines both on layer‑importance classification and Pareto‑optimal inference acceleration, and reliably identifies harmful, redundant, and critical transformer layers.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.