Steer Only When Necessary: Factorized Intervention Necessity for LLM Safety
Abstract
Activation steering provides a lightweight approach to improving large language model safety at inference time, but unnecessary intervention can cause over-refusal and degrade generation quality. These costs extend beyond benign requests: a model may already respond safely to a harmful request, making additional steering unnecessary. This motivates identifying such cases before generation while accounting for the risk of cancelling necessary interventions. To address this challenge, we introduce **Factorized Intervention Necessity (FIN)**, which separates semantic eligibility, based on hazard relevance and harmful intent, from untreated vulnerability, indicating whether the target model would respond unsafely without steering. FIN estimates vulnerability from prompt representations using offline supervision from unsteered responses. Eligible requests receive steering by default; low predicted vulnerability permits cancellation only under a rule certified on an independent held-out set. Certification assesses the probability of cancellation among eligible requests with unsafe unsteered responses. Across three primary models and diverse safety and utility benchmarks, FIN reduces unsafe leakage relative to unsteered generation while largely preserving general capabilities. Compared with the evaluated selective and adaptive baselines, FIN achieves the highest average benign compliance. Integrating FIN with multiple steering methods further reduces average quality degradation on harmful requests that the unsteered models already handle safely.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.