THE LOW-BIT MEMORY LAW: PRECISION BUYS PROPAGATION, DIMENSION BUYS RETRIEVAL, AND GATING MUST BE STABLE
Abstract
Recurrent-state architectures (linear RNNs, SSMs, Gated DeltaNet, Kimi Delta Attention) replace the KV cache of most layers with a fixed-size state, making state quantization a primary deployment bottleneck: recent work reports uniform INT8 already degrades complex reasoning and INT4 collapses to near zero. No quantitative law, however, says how long a b-bit recurrent state can remember—expressivity analyses ask whether such states can compute a function, and deployment heuristics give no account of why persistence sets risk or when the heuristic breaks. We supply the law. Under finite precision, memory is governed by three resources with closed-form bounds: (i) precision buys propagation—reach is k_max = (2√3 / eC) * 2^b steps, each bit doubling it, at a unique self-consistent forgetting rate α_*(b); (ii) dimension buys retrieval—read accuracy saturates beyond b_sat(α) = 1/2 log2( C² / (3(1−α²)) ) bits, above which 8 bits is indistinguishable from fp32 at every load while 4 bits suffices only up to M ≈ d_k/4 (the retrieval bit requirement grows with load); (iii) gating noise is exponentially amplified—one standard deviation of gate-logit noise costs 0.72 state bits. A fourth result repairs current practice: the correct per-channel persistence measure is P = 1/(1 − E[a_t²]), and the geometric-mean version used by DAMP (Zhang et al., 2026) probably underestimates risk (Jensen), by 2.7–5.7× under the occasional-reset gates that selective architectures learn; correcting it removes 99.7 ± 0.6% of the excess risk of Top-K protection across five independently trained models at an identical bit budget. That gain is a statement about gate variability, not a universal one: where the gates are nearly deterministic the correction is inert. The resulting bit-requirement log-law b_req(P) = 1/2 log2(C² P /3) explains why protected channels must be FP16, why INT4 variants cannot be repaired by layout changes, and why DAMP’s Hadamard transform works (it lowers C by up to Δb ≈ 1.24 bits on heavy-tailed blocks). Beyond closed-form verification we measure a real pretrained Mamba-1 (130M, 1536 channels): its channel statistics transfer (23× energy concentration above uniform; material damage from 4–5 bits, INT8 costing 0.12% perplexity) while the ranking correction does not, because that model’s gates are nearly deterministic—mask-to-mask variation (0.44 PPL s.d.) exceeds the correction (0.19 PPL)—so its scope is gates with non-negligible Var(ln a). All results are reproducible from released scripts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.