acceptodds
Under review as a conference paper at ICLR 2027

RefusalTrace: Localizing and Validating the Internal Refusal Pathway of LLMs with Sparse Autoencoder Features

Abstract

Safety evaluation of large language models is still largely behavioral: it records what a model outputs, not where refusal is computed or why it can be bypassed. We introduce RefusalTrace, a mechanistic framework for tracing refusal through sets of Sparse Autoencoder (SAE) features. Guided by layer-wise probing, we distinguish three processing regions along model depth: risk encoding, safety integration, and refusal execution. Within each region, we localize a representative feature set using activation-based ranking and set-level ablation under a shared refusal/compliance objective, with a general-capability control that bounds set size. We evaluate the resulting account through three-condition activation comparisons, graded suppression against activation-matched controls, and bidirectional mediation. Across Llama-3.1-8B-Instruct and Gemma-2-9B-IT, all three feature sets respond most strongly to harmful inputs. Under successful jailbreaks, risk-encoding signals are comparatively preserved while attenuation increases through safety integration to refusal execution. Interventions further identify safety integration as a key mediator between upstream risk representations and refusal execution. The localized features also support response-free harmfulness detection, illustrating their utility for safety evaluation. Code is available at https://anonymous.4open.science/r/RefusalTrace-34C1.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.