acceptodds
Under review as a conference paper at ICLR 2027

K-Shield: Improving the Safety-Utility Trade-off in Refusal Steering with Local Null-Space Constraints

Abstract

Activation steering offers a promising approach to enhancing the safety of large language models without updating their weights, but inducing refusal on malicious prompts while preserving benign utility remains challenging. Existing null-space methods address this by excluding directions associated with benign activations from the steering map's inputs. However, a direction important for benign behavior in one region of the activation space may enable malicious refusal elsewhere, so protecting it globally can hinder effective steering. We characterize the resulting refusal target reconstruction error and show how local null-space constraints can reduce it. Building on this insight, we propose K-Shield, which uses optimal transport to learn local subspaces and routes inputs to region-specific steering maps and refusal targets. Across three language models, seven jailbreak attacks, and four utility benchmarks, K-Shield consistently improves defense success and aggregate utility over global null-space constrained steering.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.