Clean Data, Unsafe Model: Certified Safety Curation for LLM Fine-Tuning
Abstract
A fine-tuned model needs safe examples of harmful prompts as much as it needs few harmful examples, and judge-based filters remove exactly those. We bring to fine-tuning data a distribution-free certificate: given any scorer and a few hundred human labels, it returns a cutoff whose kept set is at most an -fraction harmful with probability , or refuses. On BeaverTails its guarantee holds and a closed form predicts its label budget. A certified set then exposes what else a filter does. At 0.5B, a judge-selected set certified at most percent harmful trains a more harmful model than the pool, because the judge scores a response by its prompt and discards the demonstrations. On an aligned 8B model the same judge, asked SAFT's question, keeps a set more harmful than its pool and is worse than no filtering, where which examples the scorer keeps decides the harm. A scorer that confuses demonstrations with harmful examples cannot both purify a set and keep them, so the property to certify is a harmful fraction within each kind of prompt. Certified that way at 0.5B, a moderation model's selection matches training on every human-safe example at a fifth of the labels. Spectral and bilevel filter scores refuse under the same certificate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.