acceptodds
Under review as a conference paper at ICLR 2027

One Layer Is Enough: Detecting Steering from Partial Activation Logs

Abstract

Activation steering modifies a language model's behavior without touching its weights or prompts, so inspecting weights and prompts cannot detect it. We ask whether steering can be detected given only a single logged layer's activations, without access to the steering vector, the injection site, any adjacent layer, or clean-run counterfactuals. Steered states almost surely have no prompt preimage (Mishra et al. 2026), so we create an auditor that inverts the logged layer back into tokens and compares reconstruction residuals with calibrated residuals from clean reference distributions. Across five open models from the Qwen, Gemma, and Llama families, the final layer alone flags every projected gradient descent (PGD) attack at a relative budget of 0.02 and up, and at 0.0085 on three models, with only 0 to 2% of held-out clean prompts flagged. All 46 judge-positive attack-induced jailbreaks across three Qwen models are detected, and the jailbreak rate first rises at budget 0.12, approximately six times the budget at which the alarm saturates. An attacker that differentiates through the detector to stay under the alarm at only the logged layer is left with at most 1.3% of its original budget and is unable to jailbreak the model in 270 attempts, while the same attacker unconstrained jailbreaks 91 of 210 times, all of which are flagged. Within the tested attack family, attacks that went undetected induced no jailbreaks. Activation logs can therefore extend auditing beyond weights and prompts to the computation performed during inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.