JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety
Abstract
Agent safety is shifting from content moderation toward preventing operational failures before tool-using agents act. However, most agent guardrails remain reactive, judging safety only from observed actions or trajectory states. This is insufficient for long-horizon tasks, where benign-looking steps may lead to delayed harm. We introduce JANUS, a foresight-oriented framework for predictive guarding from partial trajectories. JANUS trains a shared policy with two tasks: anticipation forecasts safety-relevant futures, while adjudication judges safety from the observed prefix and anticipated future. We further propose CoAA-RL, which couples the two tasks through reward. Adjudication receives verifiable safety rewards, whereas anticipation is optimized for both future-summary fidelity and downstream utility. JANUS also synthesizes diverse trajectories through multi-agent simulation. Across four agent-safety benchmarks, the resulting model, VANGUARD, improves average protection by 15.9 percentage points over baseline guards while increasing benign task completion by 5.1 percentage points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.