acceptodds
Under review as a conference paper at ICLR 2027

HOW RELIABLY DOES REPLAY REPRODUCE A TOOL-USING AGENT’S HARMFUL OUTCOME?

Abstract

Counterfactual replay is a common instrument in replay-based attribution for tying a tool-using language agent's harmful action to a span of its context: neutralize a candidate span, re-execute, and read the outcome. Whether that verdict is a value or a draw, and how concentrated the draw is, has not been measured. We report a census. On the AgentDojo workspace suite with a stock prompt-injection attack and gpt-4o-mini, all 120 configurations are executed 20 times in each of two decoding conditions, 4,800 executions with no failed runs. The benchmark-level attack success rate is 0.563 under greedy decoding and 0.566 at temperature 1, against a round-to-round standard deviation of 0.019 and 0.026 – steadiness that is a property of averaging rather than evidence that an individual outcome repeats. In each condition a single replay of an observed harmful execution returns the opposite verdict 0.087 of the time (95% CI [0.053, 0.124]) under greedy decoding and 0.141 ([0.091, 0.189]) at temperature 1, and 0.306 ([0.183, 0.389]) and 0.404 ([0.292, 0.534]) of harmful configurations have an observed harm frequency strictly inside (0.1, 0.9) over 20 identical repeats. An aggregate error bar therefore does not bound the per-incident quantity, and on this census the two move independently. Trajectory comparison localizes the first divergence to a model output in 91 of 91 divergent pairs with byte-identical payloads, the population is far more polarized than a common-rate model allows, and that excess appears again on a second AgentDojo task suite, on a benchmark with no attacker, and on an independently implemented attacked benchmark whose checker, attack and decoding condition are none of ours. Executing the counterfactual side as well, over 6,120 preregistered executions, four interventions on the injected payload leave no harmful outcome on any contested configuration, while deleting count-matched spans that never touch it moves a configuration's harm probability almost as far – a noise-corrected root-mean-square of 0.44 [0.32, 0.54] against 0.57 [0.48, 0.65]. Sensitivity of that order recurs on a second task suite whose nine eligible configurations, from six user tasks, were selected independently: the ratio there is 1.10 [0.79, 1.75], passing a point rule fixed in advance at [0.20, 1.40], with the direction reversed – harm rises on seven of nine, each to 20 of 20, where it split eleven to fourteen here. Unsigned change alone therefore does not identify the injected source. This sensitivity is distinct from failure of signed candidate ranking and persists regardless of the replay budget used to estimate it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.