Cache Repair in Hybrid LLMs Has a History
Abstract
In recurrent-attention hybrid language models, updating the same K/V locations can produce different repairs depending on which tokens execute before each update. We study this distinction in position-independent caching (PIC), holding initialization and update locations fixed while executing short predecessor windows. The extra tokens' fresh K/V is neither read during repair nor stored in the final cache. On Qwen3.5-9B/RULER with zero-initialized recurrent matrices, increasing execution from 10.0% to 31.4% of context improves macro score over selected-only scatter by 20.9 percentage points. Installing the resulting K/V alongside scatter's recurrent and convolutional query-start states retains 18.2 points. Both effects hold with last-chunk initialization and replicate under zero initialization on 650 disjoint RULER requests. The control-study gain over naive reuse is smaller (+13.2 points), with an adjusted interval including zero. Additional history harms some OLMo tasks and reduces update coverage under fixed execution budgets. The findings distinguish what a repair stores from the execution needed to produce useful K/V, including at a fixed query-start state; they do not establish a universally beneficial history or an optimal budget allocation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.