BlockFAMQ: Realized-Residual Curvature for Complete-Checkpoint Mamba Quantization
Abstract
Mixed-precision quantization must decide which realized rounding residuals receive scarce high-precision bytes, but diagonal sensitivity scores discard the signed coordinate interactions that distinguish deterministic candidates. We introduce BlockFAMQ, a matrix-free allocator that scores squared per-token gradient–residual products and solves the resulting integer-byte multiple-choice problem under a complete-checkpoint budget. Across Mamba-1 130M, 370M, 1.4B, and 2.8B, the mean BlockFAMQ-minus-diagonal WikiText-2 NLL gap contracts from −0.218 at 2.5 effective bits to −0.026 at 4.0 bits and −0.009 at 4.5 bits. At 3.5 bits, BlockFAMQ improves six-task accuracy by 2.08 points and recovers 65.7–80.1% of the diagonal quantization gap; an activation-weighted reconstruction baseline narrows but does not remove the advantage. At 3.0 bits, direct candidate KL improves on BlockFAMQ by 0.038 nats across 1.4B and 2.8B, while a feasibility-anchored block-to-KL shortlist recovers 0.033 nats at 1.62× the block-scoring cost on average. Precision shifts toward x_proj and in_proj, whose realized residuals have the largest cross-term tails, and the block effect is larger on Mamba-1 1.4B than on matched-scale Pythia-1.4B. The direction persists on Mamba-2 2.7B and Falcon-Mamba 7B; packed W4A8 composition retains 96.8–97.7% of Quamba2 decode throughput while lowering NLL. These results support realized block curvature as a budget-dependent allocation correction, with direct KL becoming more useful when aggressive quantization leaves the local quadratic regime.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.