Can first-order memory repair compressed attention?
Abstract
Compressing attention by sharing key and value (KV) heads reduces KV-cache storage, but can hurt predictions. We address this loss of predictive capacity through a repair mechanism called **First-order Memory** (FoM). Specifically, FoM adds a compact recurrent state that summarizes causal features with storage independent of context length. We then study how well FoM repairs compressed attention, and assess whether it can also outperform stateless repair. Across a controlled byte-level study (one UTF-8 byte per token) and studies using Pythia-class models (up to 1B), FoM recovers prediction quality, often even outdoing comparably sized stateless MLP repair. In a held-out 410M MQA study, FoM lowers mean cross-entropy by 0.02235 nats per token over the selected MLP across five new seeds; this strong advantage is, however, not universal, and in some cases a strong MLP nearly matches mean performance. Overall, the experiments are encouraging and suggest that FoM can provide valuable repair, possibly at an additional execution cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.