acceptodds
Under review as a conference paper at ICLR 2027

Clean Data, Unsafe Model: Certified Safety Curation for LLM Fine-Tuning

Abstract

A fine-tuned model needs safe examples of harmful prompts as much as it needs few harmful examples, and judge-based filters remove exactly those. We bring to fine-tuning data a distribution-free certificate: given any scorer and a few hundred human labels, it returns a cutoff whose kept set is at most an -fraction harmful with probability , or refuses. On BeaverTails its guarantee holds and a closed form predicts its label budget. A certified set then exposes what else a filter does. At 0.5B, a judge-selected set certified at most percent harmful trains a more harmful model than the pool, because the judge scores a response by its prompt and discards the demonstrations. On an aligned 8B model the same judge, asked SAFT's question, keeps a set more harmful than its pool and is worse than no filtering, where which examples the scorer keeps decides the harm. A scorer that confuses demonstrations with harmful examples cannot both purify a set and keep them, so the property to certify is a harmful fraction within each kind of prompt. Certified that way at 0.5B, a moderation model's selection matches training on every human-safe example at a fifth of the labels. Spectral and bilevel filter scores refuse under the same certificate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.