SCAR: Parallel Span Cross-Attention Repair for Context Deletion in LLMs
Abstract
Deleting a span from a KV-cached LLM context leaves the surviving suffix encoded with content that no longer exists. Re-prefilling the suffix repairs it exactly, but the cost grows with the suffix rather than with the deletion, and selective recomputation still pushes every recomputed token through all backbone layers in sequence. We observe that each layer's cache already holds the information for estimating that layer's repair: the stale suffix states and the deleted-span states. SCAR (Span Cross-Attention Repair) exploits this with one small corrector per layer that cross-attends from the suffix to the deleted span and predicts a residual update to the suffix keys and values. The correctors read only the cache, run no backbone layer, and depend on no other corrector, so all of them run in parallel as a depth-1 repair and leave a standard decoding cache. Across seven task categories on Qwen3-8B and Llama-3.1-8B, SCAR exceeds 15%-budget CacheBlend-style selective recomputation in answer agreement with exact recomputation (consistency) on 13 of 14 backbone-category cells. On four timed Qwen settings, it runs 2.9-4.8 faster than exact recomputation and 18-33% below that budget's modeled latency. On conflicting-fact retrieval, SCAR achieves near-exact consistency where selective recomputation falls short even at 75% on Qwen and 50% on Llama; on Qwen GSM graphs, variable tracking, and conversation, it exceeds training-pool-matched KVEraser by 24-48 percentage points. At fixed 32k context with 512 deleted tokens, SCAR beats exact recomputation at every tested suffix length of at least 4k, reaching a 4.2 speedup at 30k. Higher recomputation budgets still recover more fidelity on several settings, at two to three times SCAR's latency. SCAR thus demonstrates that deletion repair can be learned directly from KV cache space and executed in a single parallel stage, decoupling repair latency from backbone depth and making long surviving suffixes cheap to edit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.