acceptodds
Under review as a conference paper at ICLR 2027

Steer Before You Act: Robust Thought-Guided Steering for Agent Behavioral Safety

Abstract

Large language model (LLM) agents perform long-horizon, complex tasks by reasoning iteratively and interacting with environments through external tool calls. Even when responding to legitimate user requests, agents may compromise behavioral safety through unauthorized or inappropriate tool calls that expose private information, disrupt services, or cause financial loss. Existing defenses often assess generated actions using auxiliary guard models, adding inference costs from verification and, when needed, regeneration. We introduce an activation-steering framework that steers the agent's internal thought representations before tool-call generation, without retraining the underlying language model or relying on auxiliary guard models. To improve steering reliability across tool-use scenarios, we propose Safety Preference Stability, which favors steering directions with consistent safety utility under direction and construction-data perturbations. We evaluate our method on Qwen3 models (4B–14B) in ToolEmu, an environment for evaluating agents on high-risk tool-use tasks. Our framework achieves up to and the unprotected agent's safety rate in in-distribution and evolving scenarios, respectively, while providing – speedups per invocation over external defenses. Moreover, our framework improves legitimate-goal completion by up to 30.4% over the strongest external guard baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.