A False Sense of Provenance: Data Laundering Defeats Training-Data Licensing Audits
Abstract
Training data is now bought and sold under contract. To check these contracts after the fact, researchers have built data-provenance audits, which look for a signal of the licensed data in a trained model. In this paper, we show that a licensee can train on the licensed dataset beyond what its license covers and defeat these audits with a single cheap operation, which we call data laundering. The licensee fine-tunes a single generator on the licensed dataset, regenerates the dataset from that generator, and trains each deployed model on the regenerated dataset. Any signal an audit looks for is either separable from the capability the licensee paid for or inseparable from it, and neither case attributes a laundered model to the licensed dataset. We test the attack against nine published state-of-the-art audits over images and text. Laundered models score at or near the level of models trained without the licensed data. Language models trained on the regenerated dataset keep 82% to 85% of the capability gain from training on the licensed dataset. Laundering also defeats six defenses, including two published audits designed with regeneration in mind.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.