acceptodds
Under review as a conference paper at ICLR 2027

Audit Horizons: How Long Can Stateful Agents Act Before Review?

Abstract

How many consecutive changes should a coding agent make before a trusted review? We introduce the audit horizon: the longest sequence of autonomous actions a fixed agent stack may take before trusted review, under a guarantee on regression risk. We then propose an audit-horizon qualification method: design data choose how deep and how many independent trajectories to run, and fresh trajectories determine the horizon through a valid risk-control test. We prove that one-shot evaluation cannot identify risk that accumulates along a trajectory. Certifying a nonzero audit horizon with high probability therefore requires persistent evidence. However, when earlier observations follow the same distribution regardless of the planned evaluation length, extending a run beyond the target horizon adds cost without strengthening its direct risk bound. On chained upgrades of Flask, Jinja2, and PyJWT with DeepSeek V4 Flash 0731, short-prefix qualification reaches the same audit horizon as full-length evaluation while reducing measured model cost by 53.6% and summed agent-environment time by 42.5%. Across a three-model policy comparison, sharing evaluation depth across packages achieves the highest observed horizon-cost score. Overall, we advocate that future safety metrics should reflect long-term behavior instead of only one-shot workloads.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.