acceptodds
Under review as a conference paper at ICLR 2027

Auditing Local Attention Fidelity as a Proxy for Generation Quality

Abstract

During inference in Transformer-based large language models, key-value (KV) cache compression is widely used to reduce memory overhead. Yet accurate local attention reconstruction need not imply reliable generation after compression. We test this assumption by separately measuring fitting-query residuals, operator errors at future full-model queries, and task scores from compressed generation. We use calibrated local merging (CALM) as a tractable test construction and compare it with eviction and controlled attention-fitting methods. A classical conditional Kullback–Leibler (KL) identity yields a local output bound; real-state measurements show that its global form is loose, while a retrospective group certificate reduces median prompt-level bound/error ratios from thousands to 6.40-7.95 at 40% retention. Across three near-4B model families, twelve LongBench tasks, 64 examples per task, and 8K/16K caps, we evaluate 27,648 predictions. On 24 matched Phi texts, CALM has lower operator error than ChunkKV throughout, yet lower task scores on ten. Allocation controls do not establish a geometry-specific task benefit, and compression before the first output reveals protocol sensitivity. Genuine-length inputs save 50.16-59.70% of all persistent cache state, but eager decoding is slower. These results delimit what local fidelity can support as a proxy without implying that fidelity and generation quality are generally unrelated.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.