PASVD: Adaptive Cross-Head Key Selection from Low-Rank KV Caches
Abstract
Large language models (LLMs) are widely used in interactive assistants, long-document understanding, and complex reasoning. As context windows grow, however, the key–value (KV) cache becomes a major memory and data-movement bottleneck. Many existing low-rank methods construct projection spaces offline from model weights or calibration data, leaving fixed projections unable to adapt to the evolving token distribution of each sequence. Online alternatives adapt during generation, but their approximate updates may compromise representation quality and introduce additional overhead. We introduce PASVD, a training- and calibration-free method that jointly factorizes pre-RoPE keys across KV heads within each sealed page and adaptively selects a page-specific rank from its singular-value spectrum. During decoding, PASVD computes historical query–key scores directly in the compressed space and reconstructs only the selected keys for RoPE and final attention computation. This selection-before-reconstruction design avoids reconstructing the full historical key cache while enabling content-adaptive compression and sparse access.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.