acceptodds
Under review as a conference paper at ICLR 2027

HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models

Abstract

Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance. Existing defenses either retrain or modify weights during or after fine-tuning, which requires control over training or trades safety against utility, or use model-agnostic guard classifiers, which are not tailored to a given checkpoint and run a separate large model on every prompt. We propose HyperSafe, a post-hoc, model-specific, and non-invasive framework that restores safety by generating a Safe Side Network (SSN) for each fine-tuned checkpoint. A hypernetwork maps layer-wise activation fingerprints, computed from a small set of calibration prompts, to the SSN parameters in a single forward pass. The SSN runs alongside the frozen fine-tuned model as a prompt-level safety classifier: harmful prompts are routed to refusal, while safe prompts are answered by the original model. HyperSafe targets safety degradation from benign fine-tuning rather than adversarial jailbreaks. The hypernetwork is trained once per backbone family; each new checkpoint then requires no gradient updates, no safety data, and no weight modification. On Qwen2-7B and LLaMA-3- 8B, HyperSafe reduces harmful response rates from 19–31% to below 1% on every held-out checkpoint (confirmed on LLaMA-3-8B by a second, independent judge), loses less than 1% task accuracy on average, and adds only 3.4% inference FLOPs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.