acceptodds
Under review as a conference paper at ICLR 2027

Geometry-Guided Layer Selection for Language Model Compression

Abstract

Which transformer blocks should be considered for removal? We study a layer-selection criterion that compares centered activation subspaces through projector distance. An analytical construction distinguishes this population statistic from tokenwise cosine, and a rank–angle decomposition makes its selections interpretable. On DeepSeek-R1-Distill-Qwen-14B, we compare geometry, Block Influence, and positional references using matched calibration and separate single-block interventions on 144 metallurgy questions. The geometry-selected set has a mean reference-token loss change of nats, compared with for Block Influence and positive changes for the early and late positional references. The geometry–Block Influence contrast is unresolved by the workload-bootstrap interval. All 15 geometric choices are reproduced by angular changes between retained rank-one subspaces at three calibration seeds. These results characterize an interpretable selection signal and its conditional-loss profile; they do not estimate joint deletion effects or free-generation accuracy. A separate historical compression-and-adaptation study provides resource context: a 33-block geometric configuration uses 28.0% fewer parameters and achieves the observed token throughput of the 48-block original. We distinguish candidate selection from validation of the resulting compact model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.