RISA: Risk-Informed Safety Activation for Long-Horizon Tool Using Agents
Abstract
Tool-using language-model agents increasingly operate over long interaction horizons, where unsafe outcomes may emerge only after a sequence of individually plausible actions. Existing safeguards often focus on the current action or compress trajectory history into a single risk decision, providing limited control over when and how strongly to intervene. We introduce RISA (Risk-Informed Safety Activation), a framework for long-horizon agent safety that combines multi-horizon future-risk prediction with adaptive representation steering. RISA models the distribution over the first future safety violation, derives risk estimates across multiple horizons, and uses these predictions to train an intervention policy that adaptively controls the strength of a safety-relevant internal representation direction. Complete trajectories, consisting of the full sequence of agent inputs, actions, and tool or environment observations, provide privileged supervision during training, while inference depends only on the observed trajectory prefix. Across multiple long-horizon agent-safety benchmarks and model families, RISA improves unsafe completion and safe completion while preserving task utility and limiting over-refusal compared with step-level, trajectory-aware, fixed-steering, and threshold-based safeguards. Additional analyses show that the learned safety direction captures behaviorally relevant structure, multi-horizon conditioning supports more selective intervention, and RISA shifts risky trajectory states toward lower predicted future risk. These results suggest that separating future-risk forecasting from adaptive internal control provides an effective approach to proactive safety intervention in long-horizon agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.