A Memory That Only Looks Like One: A Hybrid Vision-Language Model Relays the Picture Rather Than Storing It
Abstract
A hybrid vision-language model carries two kinds of memory. Its attention layers keep a growing cache: every token the model reads, including every image patch, is stored as a key-value entry that later tokens can look up. Its recurrent layers keep a single state of fixed size, and every new token overwrites part of that state. Prior work proved that a fixed-size state must eventually lose information, but not when a given model forgets, what it keeps while forgetting, or what breaks downstream. We answer these three questions for Zamba2-VL-7B, whose 81 decoder layers all carry a Mamba2 mixer while 13 also carry attention. First, just after reading an image the state holds the gist, not the details: a linear probe reads which object was present at on the best layer against chance, but can barely read how many objects there were or where they sat. Second, how fast the state forgets can be computed from the weights alone: each head has a half-life of tokens, with medians of , and at three scales, nothing fitted. Third, and centrally, the memory a probe appears to find at distance is mostly not memory. With write amplitude held fixed, heads with a -token half-life still decode identity at sixty-four tokens after the image, where their own parameters retain of what was written. Evicting the image's key-value entries, without touching the recurrent state, drops the fast tercile from to at thirty-two tokens, below chance, while evicting an equal amount of pre-image text moves the readout the other way; at matched amplitude, eviction costs fast heads times what it costs slow ones. Causally, exchanging all 81 recurrent states between runs that saw different images leaves of answers following the host's own image; exchanging seven of thirteen attention caches moves to the donor's. The recurrent channel relays what attention holds; it does not store the picture. The redundancy premise behind visual token pruning survives down to a quarter of the entries, but the cost of pruning harder roughly doubles with distance from the image.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.