DynSVD: Dynamic SVD via Token-Adaptive Basis Selection
Abstract
Existing Singular Value Decomposition (SVD)-based compression methods construct static low-rank bases that are fixed across all inputs, even though the contribution of each basis depends jointly on its singular value and the input token. As a result, singular directions with small singular values, which can be highly activated for specific inputs, are permanently discarded by static truncation. To address this issue, we propose DynSVD, a training-free, token-adaptive inference framework that selects bases dynamically on top of static SVD. We prove that dynamic selection is optimal for each token and that its error adds to the static truncation error without cross terms, which lets DynSVD complement static SVD methods. Experiments show that, at the same computation budget, DynSVD lowers perplexity by 42% on average relative to static SVD. Within a fixed GPU memory budget, one stored DynSVD model offers a flexible, continuous range of operating points through a single threshold, and with custom Triton kernels it yields up to higher decoding throughput at the same perplexity and up to 27% lower perplexity at the same throughput than static SVD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.