acceptodds
Under review as a conference paper at ICLR 2027

Decoupling Coordinate Permutations from Low-Dimensional Quantiles in LLMs

Abstract

Post-training weight modeling and quantization schemes predominantly treat neural network tensors as monolithic coordinate arrays, simultaneously optimizing value discretization and positional assignment. In this work, we investigate a fundamental question: How much information truly resides within the scalar values of a Large Language Model (LLM) weight tensor when strictly decoupled from coordinates? We formalize an orthogonal decomposition of any weight matrix W ∈ Rm×n into a sorted value spectrum s = sort(vec(W)) and a spatial permutation π ∈ SN . Stripping away positional entropy reveals an unexpected empirical regularity: the position-agnostic value spectrum of trained LLM tensors is universally governed by an ultra-low-dimensional continuum. Across diverse open-weight architectures (Meta-Llama-3.2, Qwen-3, Gemma-4) and across every internal layer type, an unconstrained 5-term Metalog quantile distribution fits the empirical spectrum with near-zero error (MAE < 0.0002, relative MSE < 0.002). This closed-form parameterization decisively outperforms classical Gaussian, extreme-value, and generalized quantile baselines. Crucially, our findings expose an extreme information-theoretic asymmetry: while the value distribution collapses to an O(1) description length of just 5 scalars (160 bits), spatial permutations retain an irreducible combinatorial entropy of ⌈log2(N !)⌉ ∼ O(N log N ) bits. We demonstrate that learned model capacity resides almost entirely within spatial permutations rather than value variance, establishing that future model compression paradigms must pivot toward structured permutation coding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.