RevokeBench: Testing Selective In-Context Revocation
Abstract
A language model can comply with an in-context request to revoke a fact on a direct query while losing valid neighboring information or later emitting the invalid value. We introduce RevokeBench, a benchmark of ERASE, REPLACE, and SUPERSEDE state transitions with paired instruction and physically edited contexts. It measures current-state accuracy, retention of valid facts, and literal forbidden-value emission as separate obligations. On 960 test episodes, none of four pre-specified aggregate non-direct contrasts in the two-checkpoint confirmatory study survives Holm correction. A separately specified ten-checkpoint panel nevertheless reveals an obligation-specific pattern: instruction-mediated revocation has 2.31 and 12.92 percentage points lower same-entity retention on IID and held-out grammar-and-domain OOD splits, respectively, and 16.47 and 25.12 points higher macro forbidden-value emission than canonical editing. Human judgments corroborate OOD retention loss. An external memory-operation bridge, though coverage-limited, further shows that current-state accuracy and retention need not move together. No-revocation controls establish broad elicitation for recovery and premise attacks, but the indirect-choice attack is largely ineffective. Qwen3-8B attention interventions and suffix recomputation provide scoped evidence about prefill sensitivity and edit-aware execution. Thus direct compliance alone does not establish selective in-context revocation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.