RAJNI: Importance-Aware, Dispersion-Scheduled Training-Free Token Pruning for Vision Transformers
Abstract
Vision Transformers incur inference cost that scales quadratically in the number of patch tokens, making token pruning an attractive training-free accelerator. We present RAJNI, built on two components: a forward-pass importance score that is exactly a token's contribution to its layer's CLS aggregate under ablation, and that we show upper-bounds (though does not guarantee) the token's effect on the model's output, under an alignment condition we validate empirically; and a closed-form scheduler that sets each layer's keep-ratio from the dispersion of that layer's own importance distribution, requiring no learned parameters, auxiliary networks, or fine-tuning. We validate the method with three kinds of evidence rather than a single accuracy table: agreement with exact gradient sensitivity across depth (up to Spearman correlation), a controlled ablation showing the score's two factors are individually insufficient, and a large-scale () analysis showing the scheduler's adaptivity compounds systematically with depth. On ImageNet-1K, RAJNI reaches an approximate throughput gain of 40% at an accuracy cost of 0.7% on ViT-Base, remaining close to the Pareto front of stronger training-free baselines (ToMe, SAINT) in the moderate-compression regime, and consistently ahead of GTP-ViT. Main-text throughput numbers use batch-aggregated dispersion statistics for GPU efficiency; we state this explicitly, since per-image dynamic sequence lengths would require custom sparse-attention kernels in a batched production deployment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.