acceptodds
Under review as a conference paper at ICLR 2027

When Cleaner Caches Break: Training Recipes Shape MLA Cache Quantizability

Abstract

Quantizing the latent cache of multi-head latent attention (MLA) changes both reconstruction error and final predictive quality. We study how these outcomes depend on the trained model and on which cache coordinates share a quantization scale. Our evaluation uses three production models and six matched AdamW/Muon checkpoint pairs across three sizes. A streaming temporal policy preserves causality by quantizing groups only when they close. Under streaming INT3, Muon has greater loss damage in every pair despite cleaner marginal statistics, while retaining better final quality. Under tokenwise INT3, Muon has worse final quality at 125M and 350M, but better quality at 600M. At INT4 Muon keeps better final quality in every pair under both policies, so the reversal is specific to the seven-level symmetric INT3 reference. The activation measurements motivate a geometry- specific interpretation: channel outliers can enlarge neighboring channels’ steps under tokenwise grouping, but not under per-channel temporal grouping. This grouping argument does not order the recipes: Muon has the lower outlier ratios yet the larger tokenwise damage at 125M and 350M. Controls that hold the closure schedule fixed while regrouping channels, and that tile the latent width evenly, preserve key recipe-dependent differences; an asymmetric INT3 quantizer, which stores one extra FP16 value per group, reduces Muon’s streaming damage by 51–77% across the six checkpoints. Achieved fit remains a possible mediator of the sensitivity differences. Token-local noise probes, full-prompt MMLU, FP8 controls, and independent engine tests through 32K complement the causal comparisons. We report all matched seeds and distinguish compression damage from final quality at the tested training budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.