acceptodds
Under review as a conference paper at ICLR 2027

GAVEL: Global Auditing via Evidence Latents for Long-Horizon Agents

Abstract

Safety evidence in long-horizon agent trajectories is often dispersed across multiple actions, with critical signals appearing sparsely, emerging only after long delays, or becoming unsafe only through their composition. Existing local guardrails focus on individual turns or short contexts, while memory-based monitors must decide what to retain before later actions reveal the significance of earlier evidence. We introduce Global Auditing via Evidence Latents (GAVEL), a two-pass architecture that separates global evidence organization from detailed record access. An Evidence Encoder learns a compact set of evidence latents from the full observed trajectory under safety supervision, while an Audit Decoder judges the original record conditioned on this learned global context. Across three safety benchmarks and four language-model backbones, GAVEL achieves the highest mean accuracy in all 12 evaluated settings, outperforming the strongest competing baseline by up to 12.6 percentage points. On LongSafety diagnostics with Qwen3-8B, GAVEL shows particularly large gains for delayed and compositional evidence, improving accuracy by 17.1 and 18.0 points over the strongest baseline in the corresponding buckets. Controlled latent interventions further show that the final prediction depends on trajectory-matched evidence context.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.