acceptodds
Under review as a conference paper at ICLR 2027

Safety Alignment for Agents via Hindsight and Foresight Distillation Tree

Abstract

LLM-driven agents can solve complex tasks, but the safety of their actions depends on the current environment state and their potential downstream consequences. An action that appears harmless in isolation may become unsafe because of prior interactions or the behavior it subsequently triggers. Existing safety mechanisms often overlook these dependencies. To address this gap, we propose the Hindsight and Foresight Distillation Tree (HFDTree), an agent safety alignment framework that evaluates actions using both prior context and downstream consequences. Hindsight examines each action in its interaction history, extracting positive and negative safety evidence by contrasting teacher and student action scores conditioned on correct and incorrect demonstrations. Foresight integrates evidence from sampled continuations to assess an action’s contribution to subsequent safety outcomes. A shared-prefix trajectory tree unifies these perspectives, propagating downstream evidence backward while preserving action order to assign turn-level credit relative to alternative branches. HFDTree reuses existing rollout groups and falls back to safe reference supervision when both local evidence and propagated credit are weak. Experiments on three agent safety benchmarks with two backbone models show that HFDTree significantly reduces the execution of unsafe actions without degrading general capabilities. Our code is available at  https://anonymous.4open.science/r/HFDTree-EDB4.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.