acceptodds
Under review as a conference paper at ICLR 2027

Forget One Sense, Preserve Another: Sense-Level Unlearning in Large Language Models

Abstract

When a name or word has multiple interpretations, an unlearning request may require a large language model (LLM) to suppress answers about one while preserving correct answers about another. Evaluating such a request requires testing the retained interpretation directly, since retention measured on other items cannot establish whether that interpretation is preserved. We formulate this task as sense-level unlearning. We introduce SensePairs, to our knowledge the first benchmark for LLM sense-level unlearning that separately probes two distinct, specified interpretations and requires their shared name or word in every question in its paired evaluation. Human evaluation measures retained-answer correctness and distinguishes refusals and degenerate outputs from well-formed, non-abstaining responses that omit the queried target fact. We propose an inference-time method SenseClamp that selects queries by interpretation and, within a frozen model, sets hidden-state coordinates along a contrast direction to layer-specific reference values fitted from retained contexts. Across three model families and two evaluation sets, we compare SenseClamp with existing unlearning approaches and identify retained-interpretation damage that other-item retention misses. This work provides a task formulation, paired benchmark, and inference-time intervention for unlearning across interpretations of a shared name or word.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.