acceptodds
Under review as a conference paper at ICLR 2027

Linguistic Loopholes in LLM Unlearning: From a 174-Language–Script Benchmark to Coverage-Aware Unlearning

Abstract

Unlearning a fact in one language does not guarantee its removal in others as changing the query or even the requested answer language can reopen seemingly forgotten knowledge—a cross-lingual loophole. The most straightforward solution to this challenge – unlearning in all languages – is neither scalable nor desirable as it amplifies damage to unrelated model capabilities. We introduce the task of language budgeted multilingual unlearning with the goal of selecting a subset of languages that maximize cross-lingual erasure. To study this task we introduce Cross-lingual Unlearning Tensor, an unlearning benchmark that spans 174 language–script pairs and 25 atomic paraphrase types to examine when forgetting generalizes across linguistic expressions of the same knowledge. We further propose COVER, which selects source languages to maximize predicted COVERage of languages receiving no forget supervision; enabling unlearning on a language budget. Surprisingly, we find naively selecting strong individual sources does not reliably compose into strong source sets motivating our development of COVER. At deployment COVER only requires benign calibration data and access to the frozen model. Across three model families and two disjoint forget sets, COVER reduces mean held-out residual access by 7.8–27.3% relative to uniform source selection. We find these gains extend beyond synthetic benchmarks to real news documents in low-resource languages settings using human translated data from the Low Resource Languages for Emergent Incidents (LORELEI) corpus.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.