acceptodds
Under review as a conference paper at ICLR 2027

JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

Abstract

Agent safety is shifting from content moderation toward preventing operational failures before tool-using agents act. However, most agent guardrails remain reactive, judging safety only from observed actions or trajectory states. This is insufficient for long-horizon tasks, where benign-looking steps may lead to delayed harm. We introduce JANUS, a foresight-oriented framework for predictive guarding from partial trajectories. JANUS trains a shared policy with two tasks: anticipation forecasts safety-relevant futures, while adjudication judges safety from the observed prefix and anticipated future. We further propose CoAA-RL, which couples the two tasks through reward. Adjudication receives verifiable safety rewards, whereas anticipation is optimized for both future-summary fidelity and downstream utility. JANUS also synthesizes diverse trajectories through multi-agent simulation. Across four agent-safety benchmarks, the resulting model, VANGUARD, improves average protection by 15.9 percentage points over baseline guards while increasing benign task completion by 5.1 percentage points.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.