The Ceiling on Violation-Free Success Can Sit Below τ-bench's Resolution
Abstract
τ-bench grades an agent by the database state it ends in, so an agent that reaches the goal through a write the user never sanctioned still scores a success. Two numbers read off a baseline arm decide, before any compliance intervention is built, whether its gain on violation-free success could ever be claimed: the ceiling, pass − vfpass, the most a success-preserving remover could add; and the resolution, the 95% half-width of a task-paired violation-free-success delta under a task-clustered bootstrap. On our 3B retail arm the ceiling is 1.16 pp and the resolution is 1.9–5.5 pp, so no success-preserving remover could post a claimable gain there. The bar is set by our comparison rather than by the benchmark: a task-paired test resolves an ideal remover's gain on most of our baseline estimates, and fewer than half clear the registered spread bar. We then run the check on a real intervention, a 1.5B violation-forecasting world model deployed training-free in the planning loop over two domains, two agent families, three scales and 11 conditions. It removes violating episodes at 3B, on an arm that succeeds less often than a do-nothing agent, and again on the second agent family. But most of what it stops the environment would have rejected anyway, the removal sits on the dose–response line that three uninformed arms fix, and its premium over an uninformed scorer at matched dose is unresolved at 3B and does not replicate on the second family; a minority of the paired contrasts meet the registered spread criterion together with a paired-interval condition added later. We propose the target arm's ceiling and resolution, then a do-nothing anchor and an uninformed control at matched dose, as the entry bar for compliance interventions, and release the harness and records in the supplementary archive.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.