acceptodds
Under review as a conference paper at ICLR 2027

TrueSlice: Zero Inference Overhead Pruning via Global Residual Stream Rotation

Abstract

To alleviate the significant memory and computational costs of large language models, numerous structural pruning methods have emerged, including residual stream width reduction. However, some residual width pruning methods introduce architectural changes and additional inference overhead. We propose TrueSlice, a calibration-based method that reduces residual width using a single global rotation derived from trace-normalized activation second moments. By retaining the original model architecture, TrueSlice can be applied to both standard transformer and modern hybrid networks, such as those in the Qwen3.5 family. Trace normalization balances contributions across the network, while the shared rotation enables slicing without additional residual rotation parameters or operations at inference. At matched total parameter reduction targets of 10-30%, TrueSlice achieves lower WikiText-2 perplexity than SliceGPT in 17 of 24 comparisons across eight Llama 3 and Qwen3 models without recovery training. On Llama-3.1-8B at 30% target reduction, TrueSlice achieves higher average task accuracy than SliceGPT after 3B recovery tokens. TrueSlice also complements attention head and MLP feature pruning: at a 50% target parameter reduction on Llama-3.1-8B, the best tested combination with Týr-the-Pruner raises average accuracy across five reasoning benchmarks from 44.86% to 48.01%, a gain of 3.15 percentage points over Týr alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.