acceptodds
Under review as a conference paper at ICLR 2027

Disentangling Verbalized Evaluation Awareness and Refusal through Neuron-Level Intervention

Abstract

As large reasoning models (LRMs) become more capable, reliable evaluation becomes essential. LRMs can sometimes recognize that they are being evaluated, a phenomenon known as evaluation awareness. This phenomenon raises concerns about whether safety behavior, particularly refusal behavior, generalizes from evaluation to deployment. Prior work manipulates verbalized evaluation awareness (VEA) through activation steering and observes higher refusal rates when VEA is present, but the mechanisms underlying this connection remain poorly understood. We identify VEA neurons by contrasting their activations in VEA and non-VEA reasoning segments. Intervening on 5,000 VEA neurons reduces the VEA rate from 12.6–51.8% to at most 1.4% across the model benchmark combinations in our main experiments. In QwQ-32B, VEA-neuron intervention also reduces refusal even in rollouts without VEA. We further show that their down-projection vectors contain refusal components, and preserving these components during intervention partially restores refusal while maintaining VEA suppression, providing a mechanistic explanation for the connection between VEA and refusal. Finally, across the OLMo 3 SFT, DPO, and RLVR checkpoints, the VEA rate on SorryBench and StrongREJECT is similar after SFT and DPO but substantially higher after RLVR, alongside an increase in the number of VEA neurons. Together, these findings provide a new way to intervene on VEA, enabling more reliable safety evaluations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.