acceptodds
Under review as a conference paper at ICLR 2027

RAS: Measuring LLM Safety Through Refusal Alignment

Abstract

Repeated safety evaluation of open-weight large language models requires generating and judging responses for many related checkpoints. We introduce (Refusal Alignment Score), a reference-relative measure of harmful-request refusal computed by the five-stage \SafeVec procedure. \SafeVec extracts layer-wise refusal directions from an aligned reference model, measures their alignment in related targets, and calibrates a 0–100 score using behavioral attack success rates (ASR) from calibration models. Subsequent target scoring uses input forward passes alone. Across 18 held-out variants from the Llama, Gemma, and Qwen families, correlates with (Pearson , Spearman ) and achieves 83.0% pairwise ordering agreement. Across five Llama prompt configurations, reference directions have mean cosine similarity . With calibration ASR already available, setup-inclusive evaluation of five targets is – faster than the specified behavioral protocol. These results support refusal alignment as an efficient signal for screening related checkpoints and prioritizing behavioral audits.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.