acceptodds
Under review as a conference paper at ICLR 2027

SCOPE: Sparse Coordinate Suppression Complements Activation Steering in Vision-Language Models

Abstract

Activation steering improves the safety of Vision-Language Models (VLMs), but addressing residual jailbreak failures through stronger steering can degrade benign utility. We find that pre-generation activation differences between prompts yielding safe and unsafe responses are concentrated in a sparse set of hidden-state coordinates. We introduce SCOPE, a training-free complement that zeroes these coordinates at one language decoder layer, only at the final prompt position during prefill. The support and layer are selected once from each unmodified base VLM and reused across steering methods and attack benchmarks. Coordinate and operation controls show that the gains depend on both support selection and zeroing. Across two VLMs, three steering methods, and three jailbreak benchmarks, SCOPE reduces attack success rates by 1.78 to 14.80 percentage points in all 18 evaluated combinations while largely preserving benign utility. At the example level, SCOPE recovers residual steering failures and exhibits combination-specific gains on prompts that neither intervention handles safely on its own.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.