acceptodds
Under review as a conference paper at ICLR 2027

Does Hand-Designed Trajectory State Improve Pre-Execution Control?

Abstract

We test whether accumulating risk across an agent’s trajectory improves prevention over evaluating each action independently. Across separately analyzed live model studies and released AgentDojo traces, the evaluated hand-designed state establishes no prevention advantage over a stateless per-action rule. Replaying frozen calibration grids raises trajectory prevention to the stateless level but not above it. The live attack studies share seven authored scenarios, and we do not evaluate learned sequence policies. A diagnostic identifies why one apparent observability failure occurred: the controller discarded a tool-name signal already available before execution. Separately, removing returned tool content causes model evaluators to misidentify injection origin in 33 of 40 blinded packets. Within these evaluations, accumulated risk does not improve prevention, whereas retained content supports incident reconstruction. The resulting design lesson is to preserve decision-relevant information before compressing it into a risk state, and to evaluate prevention separately from reconstruction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.