Decoupling Coordinate Permutations from Low-Dimensional Quantiles in LLMs
Abstract
Post-training weight modeling and quantization schemes predominantly treat neural network tensors as monolithic coordinate arrays, simultaneously optimizing value discretization and positional assignment. In this work, we investigate a fundamental question: How much information truly resides within the scalar values of a Large Language Model (LLM) weight tensor when strictly decoupled from coordinates? We formalize an orthogonal decomposition of any weight matrix W ∈ Rm×n into a sorted value spectrum s = sort(vec(W)) and a spatial permutation π ∈ SN . Stripping away positional entropy reveals an unexpected empirical regularity: the position-agnostic value spectrum of trained LLM tensors is universally governed by an ultra-low-dimensional continuum. Across diverse open-weight architectures (Meta-Llama-3.2, Qwen-3, Gemma-4) and across every internal layer type, an unconstrained 5-term Metalog quantile distribution fits the empirical spectrum with near-zero error (MAE < 0.0002, relative MSE < 0.002). This closed-form parameterization decisively outperforms classical Gaussian, extreme-value, and generalized quantile baselines. Crucially, our findings expose an extreme information-theoretic asymmetry: while the value distribution collapses to an O(1) description length of just 5 scalars (160 bits), spatial permutations retain an irreducible combinatorial entropy of ⌈log2(N !)⌉ ∼ O(N log N ) bits. We demonstrate that learned model capacity resides almost entirely within spatial permutations rather than value variance, establishing that future model compression paradigms must pivot toward structured permutation coding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.