Auditing KV-Cache Retention with Saved Attention Outputs
Abstract
Can saved attention outputs certify a retention decision after the original KV cache is released? We study the evidence needed for this task. For all group-retention outputs at one query, we attain the minimum native observation count at every Value width. We also derive finite-error stability bounds and an ill-conditioned family. With a retained reconstruction, an exact normalizer-residual identity constrains one shared physical Key-Value error using causal output records. It yields grid-free convex certificates for later queries, masks, and output directions. These inputs may be selected after recording. Known appended rows enter without additional uncertainty. Checking fixed coefficients is linear in cache size at fixed record count. A sensitivity condition and same-code source witnesses establish strict information gain. Frozen-model checks demonstrate this gain and expose a nonlinear-error-dominated regime. A matched-byte control favors finer Value quantization over saved outputs on the tested short prefixes. Thus the certificate quantifies the value of evidence for a declared workload; it also identifies when another storage allocation is preferable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.