ReCue: Reinstatement through Query-Cued Rehearsal for Fast Weight Models
Abstract
Transformers retain past key–value pairs explicitly, enabling strong associative retrieval but requiring a growing KV cache. Fast weight models offer a more efficient alternative by compressing the history into a bounded state, but this compression can degrade individual key–value associations through interference. Rehearsing past associations may recover them, yet fast weight updates are non-local and can perturb other associations. We introduce ReCue: Reinstatement through Query-Cued Rehearsal, a training-free method that preserves associations along underrepresented key directions in a bounded cache and selectively rehearses them using the current query. ReCue jointly rehearses the cued associations through ridge regression, applies the resulting correction only to the current read, and immediately restores the original state. On pretrained LaCT and DeltaNet models, ReCue improves associative recall on RULER and improves average performance on LongBench over the base models and query-independent rehearsal. These results show that query-cued rehearsal can recover information lost through fast weight compression.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.