acceptodds
Under review as a conference paper at ICLR 2027

Select, Then Train: Efficient Transformer Adaptation Via Neuron Selection

Abstract

As pretrained transformers continue to scale, reducing both training and inference cost is critical for efficient downstream adaptation. On one hand, parameter-efficient fine-tuning methods reduce training overhead, but they generally do not reduce inference cost. On the other, most pruning methods focus on preserving task-agnostic pretrained behavior rather than post-adaptation performance. We study structured pruning (i.e., pruning neurons, not just weights) for downstream adaptation and show that simple activation statistics measured on a small unlabeled target-task calibration set provide an effective signal for identifying task-relevant neurons. In particular, we observe strong layer-wise complementarity between activation magnitude and activation frequency. Based on this observation, we propose Select, Then Train (STT), a lightweight select–then–train framework that removes under-utilized Feedforward Neural Network (FFN) neurons prior to adaptation, producing compact subnetworks optimized for task-specific performance and efficient inference. Across multiple Vision Transformers and Large Language Models, STT achieves better accuracy–efficiency trade-offs than magnitude-based and weight–activation pruning baselines, yielding average inference speedups of approximately 1.2× on LLMs and 1.7× on Vision Transformers across diverse configurations. These results demonstrate that activation-driven structured pruning is an effective and practical approach for task-specialized transformer adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.