acceptodds
Under review as a conference paper at ICLR 2027

Compression Hacking: Towards a Unified Informatics and Geometric Perspective on Intrinsic Evaluation of Large Language Models

Abstract

Recently, the concept of “compression as intelligence” has provided a novel informatics metric perspective for large language models (LLMs), emphasizing that highly structured representations signify the intelligence performance of LLMs. However, from a geometric standpoint, the representation space of highly compressed LLMs tends to degenerate into a highly anisotropic state, which hinders the LLM's ability to comprehend instructions and directly impacts its performance. We found this compression-anisotropy synchronicity is essentially the “Compression Hacking” in LLMs representations, where noise-dominated directions tend to create the illusion of high compression rates by sacrificing spatial uniformity. Based on this, we propose three refined compression metrics by incorporating geometric distortion analysis and integrate them into an intrinsic-evaluation framework. The refined metrics exhibit strong alignment with the LLM's comprehensive capabilities, achieving Spearman correlation coefficients above 0.9, significantly outperforming both the original compression and other internal structure-based metrics. This confirms that compression hacking substantially enhances the information-theoretic perspective of LLMs by incorporating geometric distortion of representations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.