acceptodds
Under review as a conference paper at ICLR 2027

Language-Aligned Training for Transferable Erasure in Multilingual LLMs

Abstract

Large language models are nowadays multilingual. This means they can process and generate text in many languages, and the same fact can be represented across languages. While this behavior is desired, it also makes it harder to unlearn data, such as outdated, private, or harmful facts. This is because even when a fact is unlearned in one language, this fact can often still be accessed through prompts in other languages. Prior work either characterizes this gap or addresses it during unlearning itself, by adapting the unlearning objective per language, restricting updates to language-agnostic layers, or projecting out shared knowledge subspaces. These approaches often require forget data in several languages and still yield non-uniform forget and retain accuracies across languages. In this work, we observe that English-only unlearning transfers more strongly to languages whose prompt representations are better aligned with their English counterparts, as measured by cross-language retrieval accuracy. Based on this observation, we propose LATTE (Language-Aligned Training for Transferable Erasure), a method that aligns the representations of the same facts across different languages already during fine-tuning. Afterwards, any standard unlearning method can be applied without modification, using only single-language forget data. Our evaluation on Llama-3.2-3B-Instruct and Qwen2.5-3B-Instruct across up to ten languages shows that LATTE learns multilingual facts as well as standard fine-tuning. We additionally show that ours can be combined with any state-of-the-art single-language unlearning method and improves cross-lingual forgetting while maintaining retain performance, both in the unlearning language and in the others. Thereby, LATTE makes a significant step towards reliable unlearning in multilingual state-of-the-art models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.