acceptodds
Under review as a conference paper at ICLR 2027

Beyond Reconstruction: Auditing Sparse Features for Physical Channel Pruning

Abstract

Sparse autoencoders (SAEs) describe activations, but structured pruning must choose physical channels. We audit whether SAE features improve that decision in Qwen3-4B. Matched SAE and TopK-PCA representations generate fixed-width masks with the same joint search; statistical selectors receive common recovery rules. At 35.53% MLP-channel removal, SAE improves seven-task mean accuracy by 0.285 points and lowers WikiText perplexity by 0.834 versus RMS across three calibration seeds. PCA has comparable accuracy (49.52% versus 49.50% for SAE). Larger SAEs trained on over one million activation positions improve reconstruction and some local risk rankings, but neither greater sparsity nor active decoder subspaces consistently beat direct-gradient deletion scores. Common mean recovery changes rankings; FLAP-score and RCPU-score adaptations exceed SAE accuracy in the measured setting. A separate larger-calibration audit favors SAE on perplexity but favors a random dictionary on task accuracy. These results identify a limited reconstruction-proxy benefit without establishing SAE-specific functional preservation or superior physical deletion decisions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.