acceptodds
Under review as a conference paper at ICLR 2027

When Can Dictionary Reuse Compress a Composable Attention Cache?

Abstract

A compressed attention prefix must save storage while remaining correct when new tokens are appended. Reusing key and value components offers one route, but it also requires preserving which component pairs occur and how much attention mass they carry. We study additive dictionaries with an occupancy matrix and give an exact contraction that avoids expanding all pairs. Under distinct combined keys, the minimum number of fixed positive product components needed for exact composition is the nonnegative rank of occupancy. A full product of unit-norm tokens has an exact state with dictionary rows, whereas every proper raw-value subset has uniform error at least on a specified query ball. Conversely, four-cycle occupancy requires four positive components even on sixteen finite query/tail tests at nonzero tolerance. A joint perturbation bound separates occupancy and dictionary errors. Fixed-dictionary query continuation verifies the predicted covariance derivative. On 540 saved-cache cases, increasing the positive component count reduces error but quickly exhausts the storage saving: exact pair indices cost less and retain exact occupancy. Ordinary clustering still achieves lower output error. The results characterize when dictionary reuse supports composable compression and which information must survive the reduction in storage.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.