acceptodds
Under review as a conference paper at ICLR 2027

GuardValue: Estimating One-Step Values and Policy Effects with Cloned Trajectory Prefixes

Abstract

Pre-execution guards decide whether a tool-using agent's proposed action runs. Comparing guarded with unguarded episodes measures a policy that may intervene repeatedly, not the value of one intervention. We introduce GuardValue, which clones a live benchmark episode to measure both from a shared history and sums one-step values over guarded-policy visits, as the performance-difference identity prescribes. In a ten-arm matched-budget audit on 278 τ²-bench tasks, type-blocking raises success by 0.264 where the reference uses no irreversible action and lowers it by 0.101 elsewhere. One silent refusal has near-zero value, while repeated refusal changes success by -0.089 per first mutating proposal in retail, +0.270 in airline and +0.227 with a second actor. With every visit measured under deterministic serving and synchronized starts, the occupancy sum equals separately measured effects task by task in four checks covering silent and noticed refusal and a budget-charged engine, and on 47 of 50 tasks without it. A fixed visit schedule leaves the registered silent test unresolved, and production-serving airline attempts disagree. For the second actor, first-point values rank the engine above silent refusal, whose effect is ten times larger. First-point values should not be reported as policy effects.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.