acceptodds
Under review as a conference paper at ICLR 2027

FCRS: Learning Fine-grained Safety Signals with Cross-layer Consistency for Refusal Steering

Abstract

Steering models to refuse malicious prompts is vital for a variety of LLM safety applications. However, achieving robust refusal activation steering remains challenging because operating on the full activation vector limits the capture of fine-grained safety signals and treating each layer hinders cross-layer consistency. To address these challenges, a novel refusal activation steering method is proposed by integrating fine-grained activation unit selection with cross-layer consistency modeling, including a fine-grained activation unit selector learning (FAUSL) module and a cross-layer consistent safety-relevant signal detector learning (CCSDL) module. The proposed method enjoys several merits. First, the proposed FAUSL module can select the most safety-relevant AUs effectively via an AU importance scoring network. Second, the CCSDL module can model long-range dependencies across layers to extract consistent safety-relevant signals. Extensive experimental results on eleven challenging benchmarks show that our proposed method significantly outperforms state-of-the-art refusal activation steering methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.