acceptodds
Under review as a conference paper at ICLR 2027

Compression Hypercube: A Unified Framework for Pruning, Quantization, and Distillation

Abstract

Model compression is critical for deploying large language models (LLMs), yet selecting an effective strategy remains largely empirical: pruning, quantization, and knowledge distillation (KD) are often studied in isolation, with limited guidance on when each method is reliable. We propose the Compression Hypercube, a structured framework that characterizes compression as an interaction between method, compression level, recovery strategy, initialization state, and data modality. We identify three key properties governing compression behavior — sequential redundancy, loss-landscape curvature, and compression-induced distributional shift — and derive local bounds linking these to compression error. Across synthetic environments and eleven real pre-trained and instruction-tuned LLMs up to 14B parameters, we observe substantial variation in outcomes: post-recovery perplexity can differ by up to depending on method choice alone. Weight-only 8-bit quantization remains near-lossless across models, while depth pruning leads to significant degradation even after recovery. Recovery behavior is regime-dependent: KD with reverse-KL or on-policy objectives dominates SFT across compression regimes, reversing earlier findings based on forward-KL only. Finally, a composite diagnostic using four base-model features (bottleneck redundancy , curvature ratio, weight kurtosis, and nominal compression level) achieves zero top-1 regret in leave-one-model-out configuration selection, providing a cheap pre-deployment protocol that requires no trial compression runs. Together, these results replace ad-hoc compression selection with a principled, theory-grounded workflow: diagnose, rank, and verify — all from a single forward/backward pass on the uncompressed model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.