acceptodds
Under review as a conference paper at ICLR 2027

DoseRank: Measured Rank Allocation for Training-Free MLA Conversion

Abstract

The key-value (KV) cache is a major memory bottleneck in large language model serving. Multi-head latent attention (MLA) reduces this footprint by caching a low-rank latent representation per token instead of separate keys and values. However, converting existing multi-head attention (MHA) and grouped-query attention (GQA) checkpoints to MLA without expensive retraining remains challenging. Prior methods allocate rank uniformly across layers or use spectral information to guide allocation, yet a layer's spectrum can be a weak predictor of its contribution to accuracy. We introduce DoseRank, an accuracy-guided rank allocator that measures each layer's sensitivity to compression using a smoothed accuracy metric on a small calibration set. At matched KV-cache sizes, DoseRank improves accuracy over uniform rank allocation by up to 5.5, 2.9, 5.6, and 7.9 percentage points on gpt-oss-120B, gpt-oss-20B, Grok-2.5, and Llama-3.1-8B, respectively. On our CARE-basis port, evaluated without recovery training, DoseRank outperforms the CARE-E spectral scheduler by up to 3.4, 2.5, 5.5, and 5.8 percentage points, respectively. Among evaluated budgets where compression leaves room for accuracy recovery, item-paired 95% confidence intervals exclude zero at most budgets for uniform and spectrum-guided allocation. Calibration on a strict subset of tasks retains most of the gains on held-out tasks, suggesting that DoseRank captures compression sensitivity that transfers beyond the calibration tasks. These results demonstrate the value of allocating rank according to measured accuracy sensitivity rather than spectral information alone. Code will be released upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.