LIVE-MLA: LATENT INFERENCE WITH VALUE ERROR FEEDBACK FOR MULTI-HEAD ATTENTION
Abstract
Most pretrained models use Multi-Head Attention (MHA) or Grouped-Query Attention (GQA). Recent methods convert these models to Multi-Head Latent Attention (MLA) through low-rank decomposition, reducing KV-cache storage without training comparable MLA models from scratch. The resulting low-dimensional bottleneck leaves reconstruction residuals that fixed low-rank mappings cannot adapt to during inference. We consider both how to correct these residuals during decoding and how to account for their effect when fitting the low-rank mappings. First, when a token leaves a small exact recent region, its latent representation and value residual are available together and can guide corrections for older compressed tokens. Second, KV reconstruction error alone is insufficient to assess compression, since keys and values matter through the attention outputs they produce. These observations lead to LIVE-MLA, which combines Residual Feedback with attention-weighted conversion. Residual Feedback fits sequence-specific value corrections from the observed pairs in closed form. Attention-weighted conversion weights each token's joint key-value reconstruction objective by the attention mass it receives. All persistent states are included in the same fixed cache budget. Experiments show that LIVE-MLA narrows the perplexity gap left by training-free conversion. Shorter recovery training further narrows the gap to the original model while using to fewer recovery-training tokens than the evaluated baseline adaptations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.