acceptodds
Under review as a conference paper at ICLR 2027

SafetyTransfer: Distilling Harmful and Harmless Discrimination for Selective Refusal in Small Language Models

Abstract

Safety alignment requires language models to refuse harmful requests while preserving useful responses to harmless ones. Transferring safety capabilities from an aligned teacher can improve compact models, but output-level distillation does not explicitly transfer the internal discrimination that determines when refusal is appropriate. We introduce SafetyTransfer, a representation-level framework that transfers harmful–harmless discriminative information from an aligned teacher to a smaller student. SafetyTransfer first selects teacher layers where this distinction is both geometrically separated and linearly accessible, then extracts a low-rank contrastive subspace rather than matching complete hidden states. Each prompt is projected onto this subspace to preserve an input-specific discriminative component. To bridge heterogeneous hidden dimensions, the component is expressed with sparse token-semantic coefficients derived from the teacher LM head and reconstructed with the corresponding student token-semantic dictionary. The student is then trained with representation alignment to these reconstructed targets together with standard response supervision. Across three teacher–student pairs and four safety datasets, SafetyTransfer achieves the highest harmful refusal rate minus harmless over-refusal rate (Gap) among the main baselines and the highest harmonic mean of harmful refusal and harmless acceptance (SR-H) in 10, while targeted removal of the transferred subspace produces larger behavioral changes than norm-matched orthogonal perturbations. The code is provided in the supplementary material.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.