EMBER: Margin-Targeted Energy Reshaping for Mitigating Exact Memorization in Autoregressive Language Models
Abstract
Autoregressive language models can reproduce protected training passages verbatim. Training-time suppression reduces this behavior, but strong suppression can also degrade ordinary generation. We propose EMBER, a margin-targeted energy reshaping objective that controls how much pressure each protected continuation receives. Using average token negative log-likelihood as energy, EMBER applies a smooth margin penalty: low-energy protected sequences receive strong repulsion, and the pressure decays as their energy rises above a specified target. The margin sets the location of this transition, while a temperature controls its width. Our analysis characterizes this gradient attenuation and connects the energy target to conditional sequence probability. Experiments on BookMIA and a controlled Wikitext-2 split cover GPT-2 with full fine-tuning and Llama-3 and Qwen-3 with LoRA. On BookMIA, EMBER reduces exact match from 34.7% to 4.4% for Llama-3-8B and from 29.4% to 1.9% for Qwen-3-8B, while keeping perplexity and MAUVE close to the unprotected models. The results demonstrate that an explicit energy target provides a simple, effective way to balance exact memorization mitigation with generation quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.