acceptodds
Under review as a conference paper at ICLR 2027

Safety Risk of Fine-Tuning Data Is Not a Property of the Sample Alone

Abstract

Large language models are safety-aligned before deployment so that they refuse harmful requests, and they are then commonly fine-tuned for downstream tasks. However, such fine-tuning can weaken safety alignment even when the training data are entirely benign. Existing defenses filter the training data by estimating the safety risk of each sample from internal model signals, such as hidden representations, sample gradients, or parameter updates. These methods typically score samples once at the initial model state and retain the resulting selection throughout fine-tuning, which implicitly assumes that the initial risk ranking remains informative as the model changes. We show that this assumption does not hold. On a controlled fine-tuning trajectory, risk rankings computed at initialization are nearly uncorrelated with those computed after two epochs. A matched intervention experiment further shows that the same groups of benign samples cause different degrees of safety degradation depending on the model state from which they are learned. Safety risk is therefore a joint property of the sample and the model state rather than an intrinsic property of the sample. We also find that rankings are rewired mainly within the first epoch. At initialization, gradient-based risk scores are dominated by a gradient component nearly orthogonal to the safety direction, and this failure persists under larger adapters, full-parameter gradients, and a short probe epoch. We therefore propose SABA, which re-evaluates sample risk at the current model state in each training round and formulates filtering as the allocation of a fixed exclusion budget across rounds. The resulting concave allocation problem reduces to a geometric schedule controlled by a single deferral parameter: no budget is spent at initialization, and exclusions shift toward later rounds with more reliable risk estimates. Under the same training budget, SABA reduces the attack success rate from 63.3% to 18.3% on CatQA and from 53.1% to 14.4% on HEx-PHI, outperforming nineteen baselines while preserving utility, and the gains carry over to a second corpus and a second model family.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.