acceptodds
Under review as a conference paper at ICLR 2027

Causal Importance Distillation for Transformer Module Pruning

Abstract

Existing adaptive inference methods for Large Language Models rely on observational heuristics, such as hidden-state similarity or activation magnitudes, to drop redundant modules. However, these correlation-based metrics often fail to capture subtle, non-linear structural computations vital for semantic accuracy. We introduce CausalGate, an intervention-guided framework for compute-efficient transformer inference. During a calibration phase, CausalGate isolates individual Attention and MLP sub-layers, zeros out their respective outputs, and measures the resulting change in the predictive distribution via the Kullback-Leibler divergence of the final logit distribution. To eliminate runtime routing overhead, this intervention-derived importance hierarchy is distilled into a global set of static, lightweight scalar gates using an Exponential Moving Average smoothing objective paired with a differentiable pairwise ranking loss. We comprehensively evaluate CausalGate on TinyLlama-1.1B across language modeling and commonsense reasoning benchmarks, with comparisons against prominent routing and module-skipping baselines. Scaling experiments on Qwen2.5-3B and Llama-3.1-8B further evaluate perplexity under increasing module removal, while hardware measurements on TinyLlama-1.1B and Qwen2.5-3B demonstrate practical inference speedups.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.