Surprise Masking: Private Data Filtering with Model Likelihoods
Abstract
Data useful for language model training can contain private or unwanted information that is at risk of being learned or memorized. We propose Surprise Masking (SM), a simple form of perplexity filtering that uses a publicly available reference model to identify unwanted spans and suppress learning from them while preserving the rest of each example. On synthetic code benchmarks, we show SM improves the privacy-utility tradeoff over standard masking methods and PII filtering. In fine-tuning experiments, SM reduces the recovery of private values from 66.3% to 2.3% and the reproduction of stylistic coding quirks from 23.6% to 1.3%, while preserving performance on the evaluated coding tasks. SM is similarly useful in pretraining. We also evaluate SM for sanitizing datasets by directly removing unwanted content. Here, SM outperforms DePA, a perplexity-based defense for detecting poisoned dead code, while specialized detectors perform better on broader credential and PII benchmarks. Across eleven settings, we find that SM's downstream success is predictable from the separation of unwanted and useful content in the distribution; particularly, we find it particularly useful with code data. Ultimately, SM is an inexpensive, adaptable method to preserve valuable examples while reducing the learning of unwanted information.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.