acceptodds
Under review as a conference paper at ICLR 2027

Steer Only When Necessary: Factorized Intervention Necessity for LLM Safety

Abstract

Activation steering provides a lightweight approach to improving large language model safety at inference time, but unnecessary intervention can cause over-refusal and degrade generation quality. These costs extend beyond benign requests: a model may already respond safely to a harmful request, making additional steering unnecessary. This motivates identifying such cases before generation while accounting for the risk of cancelling necessary interventions. To address this challenge, we introduce **Factorized Intervention Necessity (FIN)**, which separates semantic eligibility, based on hazard relevance and harmful intent, from untreated vulnerability, indicating whether the target model would respond unsafely without steering. FIN estimates vulnerability from prompt representations using offline supervision from unsteered responses. Eligible requests receive steering by default; low predicted vulnerability permits cancellation only under a rule certified on an independent held-out set. Certification assesses the probability of cancellation among eligible requests with unsafe unsteered responses. Across three primary models and diverse safety and utility benchmarks, FIN reduces unsafe leakage relative to unsteered generation while largely preserving general capabilities. Compared with the evaluated selective and adaptive baselines, FIN achieves the highest average benign compliance. Integrating FIN with multiple steering methods further reduces average quality degradation on harmful requests that the unsteered models already handle safely.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.