Class-Specific Pruning of Vision Transformers via Activation Calibration
Abstract
Existing class-specific pruning methods compress models for target classes but primarily evaluate neuron importance using activation magnitudes, treating all tokens equally. Our analysis reveals that tokens with high activation responses are not concentrated on target foreground objects; instead, many lie in background regions unrelated to the target classes, interfering with the identification of category-critical neurons. Further analysis shows that using frequency stability and attention weights to calibrate the activation contributions of different tokens can mitigate interference from irrelevant background information in neuron importance evaluation. Based on these findings, we propose Activation-Calibrated Pruning (AC-Prune), which identifies category-critical neurons based on the calibrated activation contributions. The selected neurons then serve as anchors to guide feed-forward network (FFN) neuron consolidation, while low-rankness is encouraged in multi-head attention (MHA) weights. This mitigates irreversible knowledge loss caused by direct pruning, yielding compact and efficient class-specific Vision Transformers (ViTs). Comprehensive experiments with five base ViTs covering three representative visual tasks on four datasets demonstrate that AC-Prune-derived ViTs retain approximately 20%–40% of the base models' parameters while outperforming them on class-specific tasks by up to 15.72% in accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.