acceptodds
Under review as a conference paper at ICLR 2027

Follow the Gradient: Tracing Safety Degradation Back to Training Data

Abstract

Fine-tuning Large Language Models (LLMs) on instruction data is known to degrade safety, even when no training examples look harmful. Because such data is benign at a surface level, content-based filters often struggle to filter it. We tackle this problem through data attribution, by developing a scalable gradient attribution method to identify fine-tuning examples whose gradients align with unsafe behavior. By building off of a low-rank attribution method and restricting backpropagation to the final four transformer blocks, our method maintains high embedding quality at a fraction of the cost, making it practical to run on large SFT datasets. First, in models exhibiting emergent misalignment, we find that gradient embeddings from our method connect the training data that induces misalignment to the harmful behaviors that emerge after fine-tuning, with AUC as high as 0.865. Motivated by this finding, we then apply our method to Tulu 3, scoring each example against a reference set of unsafe responses and removing the 20% of examples most aligned with them, to create SafeTulu. Fine-tuning on SafeTulu reduces GCG attack success rate by up to 22% across Qwen, Llama, and Olmo, outperforming an LLM-as-a-judge. Across multiple capability evaluations, SafeTulu mostly preserves capability relative to training on the full data. We also identify possible sources of benign-looking safety degrading data, such as tool-execution failures, based on the data our gradient filtering removed. By making gradient attribution more scalable while preserving its safety signal, our method allows for producing safer fine-tuning datasets on the order of hundreds of thousands of examples.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.