The Precision Cliff: An Information-Theoretic Lower Bound on Quantization for Iterative Compositional Retrieval
Abstract
Low-bit quantization degrades multi-hop reasoning faster than single-hop recall, yet no theory has said why deeper reasoning needs more bits or how many more. We answer both with a single closed-form information-theoretic floor. For any post-training compression scheme whose error placement is independent of the stored content and whose distortion is deterministic—conditions met by GPTQ, AQLM, AWQ, and QuIP#—the minimum bits per entry for reliable -hop retrieval is with , and an explicit construction saturates this floor to within bits. A direct corollary is the separation law: deep reasoning demands twice the bit-rate of single-hop recall at capacity saturation (predicted for ; measured ). Capacity utilization transfers this bits-per-entry floor to bits-per-parameter: a synthetic sweep traces a monotone depth penalty in (Spearman , ), and four LLM families (Llama-2-7B, Qwen3-8B, Phi-4-mini, InternLM3-8B) sit on the same -curve at production scale with wall premiums of –—already forcing a – upward correction to any quantization floor calibrated at . A chain-given control (accuracy read-aloud vs. sequential lookup at 3 bits) isolates iterative retrieval as the causal mechanism, ruling out attention-routing and generation-noise alternatives.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.