acceptodds
Under review as a conference paper at ICLR 2027

Learning to Align Where Risk Emerges for LLM Agents

Abstract

LLM agents solve long-horizon tasks through multi-turn interaction and tool use, but also introduce safety risks such as privacy leakage and unauthorized tool execution. Runtime guardrails can mitigate these risks, yet if bypassed, the underlying model may still lack a reliable basis for safe decisions, motivating intrinsic safety alignment. However, trajectory-level alignment provides sparse supervision, obscuring where to intervene for safe execution. Step-level alignment offers denser supervision, but localizing steps that steer trajectories toward unsafe behavior typically relies on stronger external models, which may be unavailable for frontier agents and incur substantial API costs and privacy risks. Surprisingly, We find that the latent risk awareness of the acting LLM Agent can instead support reliable risk localization. We therefore propose AgentFalcon, which quantifies each sentence's information gain toward an unsafe judgment elicited from the acting LLM Agent and residualizes it against the corresponding safety information gain to identify sentences that drive unsafe behavior. It then intervenes at the earliest risk position, elicits safety reflection, and pairs the regenerated safe continuation with the original continuation for preference optimization and repeat this process to improve risk coverage and preference-data efficiency. Experiments across four LLM backbones show that AgentFalcon reduces unsafe behavior while preserving benign-task performance. On Qwen3.5-9B, it reduces the attack success rate on AgentLAB from 53.3% to 8.7% and attains a 4.1% harmful-task score and 96.8% refusal rate on AgentHarm. It also improves DeepPlanning accuracy from 15.4% to 16.6% and LiveClawBench from 25.2% to 28.6% which indicate stronger capabilities in long-horizon task planning and cross-service task execution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.