acceptodds
Under review as a conference paper at ICLR 2027

Mapping Entropy: A Distributional Information-Theoretic Probe for LLM Manifold Structure

Abstract

We propose mapping entropy, a distributional, information-theoretic probe for comparing LLM representations. Each model's manifold is partitioned by a Vector Quantization (VQ) bottleneck; aligning two partitions via a shared probing corpus yields per-code conditional entropies whose distribution exposes where representations align and diverge. Scalar similarity metrics average away this per-code structure. Experiments across Qwen3-14B (sixteen depths), the Qwen3 family (0.6B–30B), and Pythia-6.9B yield three findings. First, per-code entropy distributions are highly heterogeneous: for distant layer pairs, 69% of codes exceed 6 bits, yet a zero-entropy shared base persists. Shuffled-token, frequency, and partitioner controls rule out artifactual explanations. Second, synthetic calibration shows Miller-Madow bias is substantial but systematic, and multi-seed validation establishes that VQ training noise dominates corpus sampling variance. Finally, mean conditional entropy predicts discrete-partition transfer accuracy (), whereas CKA does not () and Procrustes shows only a weak effect (). A continuous-space transfer task gives the complementary picture: CKA predicts it (), mapping entropy does not ().

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.