EinSort: Sorting is All We Need for Tensorizing LLM
Abstract
Tensor networks provide efficient representations for compressing large neural networks, but identifying useful low-rank structure in foundation-model tensors remains challenging. We propose Einstein sorted sum (EinSort), an adaptive tensorization method that exposes low-rank structure through reversible index ordering before tensor decomposition. We prove that, under stated regularity assumptions, row-wise sorting of i.i.d. entries converges to a rank-one quantile template with relative Frobenius error ; sliced or shared permutations then limit metadata overhead in practical settings. We evaluate EinSort in two deployment regimes: KV-cache compression for language models and weight compression for vision-language and vision-language-action models. Across WikiText-2 language modeling, GSM8K mathematical reasoning, TextVQA visual reasoning, and LIBERO robot manipulation, EinSort consistently improves downstream quality relative to unsorted tensor decomposition and the practical low-rank baselines evaluated in each setting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.