DoseRank: Measured Rank Allocation for Training-Free MLA Conversion
Abstract
The key-value (KV) cache is a major memory bottleneck in large language model serving. Multi-head latent attention (MLA) reduces this footprint by caching a low-rank latent representation per token instead of separate keys and values. However, converting existing multi-head attention (MHA) and grouped-query attention (GQA) checkpoints to MLA without expensive retraining remains challenging. Prior methods allocate rank uniformly across layers or use spectral information to guide allocation, yet a layer's spectrum can be a weak predictor of its contribution to accuracy. We introduce DoseRank, an accuracy-guided rank allocator that measures each layer's sensitivity to compression using a smoothed accuracy metric on a small calibration set. At matched KV-cache sizes, DoseRank improves accuracy over uniform rank allocation by up to 5.5, 2.9, 5.6, and 7.9 percentage points on gpt-oss-120B, gpt-oss-20B, Grok-2.5, and Llama-3.1-8B, respectively. On our CARE-basis port, evaluated without recovery training, DoseRank outperforms the CARE-E spectral scheduler by up to 3.4, 2.5, 5.5, and 5.8 percentage points, respectively. Among evaluated budgets where compression leaves room for accuracy recovery, item-paired 95% confidence intervals exclude zero at most budgets for uniform and spectrum-guided allocation. Calibration on a strict subset of tasks retains most of the gains on held-out tasks, suggesting that DoseRank captures compression sensitivity that transfers beyond the calibration tasks. These results demonstrate the value of allocating rank according to measured accuracy sensitivity rather than spectral information alone. Code will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.