acceptodds
Under review as a conference paper at ICLR 2027

One Patch Fits All: Optimizing Universal Safety Patches for Fine-tuned LLMs

Abstract

The fine-tuning-as-a-service paradigm allows users to customize large language models (LLMs) with their own data, but also exposes a security risk: adversaries can craft fine-tuning datasets that erode the model's safety alignment. Existing defenses either require per-request computation that scales linearly with the volume of fine-tuning jobs, or add a fixed safety vector via task arithmetic, which assumes that a vector extracted at one checkpoint has the same effect on any other. This translation-invariance assumption does not hold in general. In this work, we propose OPUS (Optimizing Patches for Universal Safety), which learns transferability by directly optimizing a universal safety patch. The framework consists of three stages: (1) it simulates a diverse set of fine-tuned checkpoints by constructing a perturbation bank through fine-tuning and interpolation, (2) it optimizes a patch so that the patched model refuses harmful queries while a KL term anchors it to the unpatched fine-tuned model to preserve utility, and (3) it deploys the pre-computed patch to incoming fine-tuned models via a single weight addition, requiring neither additional training nor user data. Extensive experiments across four models, four tasks, and three out-of-distribution attack sources demonstrate that OPUS achieves the lowest average harmful score in every setting while keeping task utility close to that of undefended fine-tuning, and incurs negligible per-request overhead.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.