Stale After Reuse: An Application Correctness Boundary and No-Regeneration Repair
Abstract
Long-context large language model (LLM) applications repeatedly revise shared context, while serving systems reuse KV states to avoid redundant computation. Prior work shows that edited information can persist in downstream KV states; we ask when this latent persistence becomes an application-record error under an executed reuse plan, how far recomputation must extend to remove it, and whether an output-visible mismatch can be corrected without another generation. In execution-proven interventions on 192 paired cases, one-gap reuse lowers exact-record accuracy from 54.17% to 45.31%. A complete chunk-aligned frontier finds every partial refresh policy below Dense marginal accuracy; complete stale-suffix invalidation restores the Dense value of 54.17%, although per-case recovery is heterogeneous. This characterization motivates correction at the output boundary, where the superseded value becomes observable. We therefore derive RevisionPatch-Exact, a deliberately conservative, no-regeneration rule that patches one unique whole-field match and otherwise abstains. Across dialogue-state and tool-argument revisions with Qwen, Mistral, and Llama models, the unchanged rule yields positive net fixes; on 384 natural MultiWOZ revisions it raises exact accuracy from 54 to 63 and changed-field accuracy from 231 to 269, with no observed exact break. The repair adds neither a model call nor prefill and, at 16k context, retains 5.85× faster median time-to-first-token and 1.61× higher throughput than Dense. These results connect stale downstream state to an execution-proven application correctness failure and expose a simple output-side operating point.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.