SeNPO: Preference Optimization in Sentence Embedding Space for Knowledge Unlearning
Abstract
Large language models memorize sensitive personal information, copyrighted material, and text that violates usage policies. Blocking a forbidden string at the output leaves that information in the weights, where later red-teaming can recover it. Unlearning fine-tunes the model so a target answer is harder to elicit, but the final token can change while an earlier layer still ranks the answer highly. We call this gap implicit evocation. A model can then look forgotten on ordinary answer metrics and still leak the target to a linear readout of its hidden states. A local residual-stream analysis gives a sufficient condition for that split and motivates a forget loss on several prefix-level semantic readouts. Semantic Negative Preference Optimization (\methodname) maps hidden states into a frozen sentence-embedding space and applies a length-normalized softplus preference loss there, together with retain-set KL. Unlike ATU's hinge loss, which provides no gradient once similarity falls below its threshold, our softplus loss remains active. Moreover, optimizes sentence-level semantic similarity, whereas token-level NPO optimizes the likelihood of the target token sequence. removes the requested facts while the model stays usable: scores on general knowledge tests and fluent writing remain close to the original model. Gradient ascent and token-level NPO, when kept from collapsing, still produce most of the answers that should have been forgotten. A matched-budget form of ATU either barely forgets or collapses, and the learning rates we tried contain no point that does both. The gap is also internal. After , a simple classifier reading intermediate hidden states recovers the forgotten answer less often. The same classifier is almost unchanged for usable gradient-ascent and NPO checkpoints. Pushing those baselines toward the same level of output forgetting still leaves that internal signal, and neighboring knowledge or fluency is lost. On fictitious-author unlearning, forgets more thoroughly than a representation-level method at similar utility. With the same retain penalty it leads token-level SimNPO on Qwen and trails it on Phi. On Llama, NPO removes more of the target answer; keeps more related knowledge and writes more fluently.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.