acceptodds
Under review as a conference paper at ICLR 2027

BARTER: Compact Dynamic Sparse Training with Standard Dense Kernels

Abstract

Dynamic sparse training (DST) aims to reduce training costs while maintaining accuracy. Existing DST methods mask out parameters but keep matrix dimensions unchanged. Training speedups therefore depend on specialized support, such as sparse kernels. We introduce BARTER, a structured DST framework that reduces training cost using standard dense kernels by keeping training tensors physically compact while dynamically reallocating capacity across layers. Unlike existing works that require computing gradients for all parameters, BARTER instead uses statistics collected only from active units during ordinary backpropagation to prune and clone attention heads and feed-forward channels under separate fixed global budgets. At each topology update, it resizes the affected parameter tensors and remaps their optimizer states, so subsequent training steps operate directly on the compact model. Training and fine-tuning experiments with Transformer-based vision models on two image classification benchmarks at densities from 0.1 to 0.7 show competitive Top-1 accuracy relative to full-tensor DST baselines. On an NVIDIA L40, BARTER achieves speedups of up to 6.29× in measured training segments and reduces peak allocated GPU memory by up to 81.3% relative to these implementations. An iPhone 13 Pro Max training benchmark additionally shows a 2.78× speedup over the full-tensor DST baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.