LaunderBench: Is Your Data Laundered and then Used in LLM Training?
Abstract
Data laundering refers to an emerging threat that rewrites documents before training a language model. This threat will hide evidence of copyrighted data use from training-data provenance detection while preserving knowledge of such source data. However, the community has yet to establish an evaluation protocol that distinguishes whether the detector is simply underpowered, the laundering succeeded, or the model failed to learn. To address this confounding, we introduce LaunderBench, a benchmark grounded in three diagnostic criteria, namely *baseline detector auditability*, *audit evasion efficacy*, *source utility viability*. Specifically, LaunderBench trains three compute-matched models from a shared checkpoint: laundered (on rewrites), unlaundered (on originals), and filler-only (on background text alone), and audits all three on the original documents. This setup allows us to separate audit evasion from detector weakness, and distinguish whether the model actually learned from the sources or merely drifted during background training. Evaluating across model scales (Pythia-1.4B to 6.9B), architectures (Pythia and GPT-Neo), 4 data corpora, and 9 rewriting methods reveals three findings: 1. rewrite-trained models consistently learn source knowledge, but the utility gap lost relative to unlaundered training widens as training proceeds; 2. the acquired source knowledge does not translate into higher detectability; 3. under strict false-positive constraints, the relative evasion rankings of rewriting methods can invert. These findings show that reliable provenance audits must evaluate detector baseline power alongside strict false-alarm budgets, which LaunderBench operationalizes to enable rigorous benchmarking of laundering-resilient data provenance audits.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.