Separating Source Efficacy from Cross-Context Transfer of Refusal Directions
Abstract
Does a frozen refusal-direction intervention work in its source context, and how does the same intervention behave in each target context? We evaluate published directional ablation using paired, context-specific baselines. On Qwen2.5-7B, the original Site A candidate has no established positive source or Passive effect under a fixed single-layer operator. Changing the extraction site to B reveals semantic harmful-compliance increases of 59.0/51.5 percentage points at source and 37.0/13.9 on Passive (JailbreakBench/HarmBench). A weak extraction can therefore hide, under the same operator, an intervention that still raises Passive compliance. A second extraction, screened only on Questioning validation, is source-effective on held-out requests, though the weaker candidate's increment is small (+6.1 points; interval excludes zero). The candidate gap remains large on low-baseline Neutral, with no clear contraction detected (D_N = -1.0 points, 95% interval [-9.4, 7.3], including zero), and shrinks on Passive, whose no-intervention rate is already 65.3%. A review of 30 requests supports these two qualitative conclusions. Academic shows automatic contraction; human review is same-signed but uncertain. Checking source efficacy is not a substitute for measuring each target wrapper.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.