CAST: Cascaded Adaptive Safety Triage for Agent Trajectories
Abstract
Safety classification of tool-use trajectories requires distinguishing unsafe behavior from legitimate actions whose safety depends on the surrounding request, permissions, and outcomes. However, trajectory-level binary labels do not explicitly reveal these safety-relevant dependencies, creating two challenges: learning a context-sensitive safety boundary from limited supervision and deciding when a prediction requires stronger review. We present CAST (Cascaded Adaptive Safety Triage), a cascaded safety monitor that addresses both. Evidence Construction and Counterfactual Editing creates evidence records and targeted label-flipping trajectory edits to expose local safety boundaries. Privileged Expert Distillation transfers evidence-informed supervision into an Expert that requires only the original trajectory at inference time. Three-Action Routing then learns to predict safe, predict unsafe, or defer difficult cases to the stronger Expert. On AgentSafety-Bench, CAST-Pro achieves 89.82% Macro-F1 and outperforms the compared guardrail and post-training baselines. CAST-Flash retains 86.19% Macro-F1 while reducing Expert calls to 28.50% and Expert-token usage to 30.79% of Pro. Across four additional benchmarks, CAST also improves average Macro-F1 over same-backbone post-training baselines. These results show the value of jointly learning context-sensitive safety judgments and selective Expert review. Code and data are available at https://anonymous.4open.science/r/CAST-ICLR2027.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.