acceptodds
Under review as a conference paper at ICLR 2027

Correction Has a Window: How Long a False Premise Stays Fixable in Tool-Using Agents

Abstract

A tool-using agent reads a record, the record is wrong, and later the truth arrives. We ask how much the timing of that correction matters. We overwrite one field in a tool result that a -bench retail agent has already read, let the agent act on it for live turns, and then delete the false record and append the true one. Across ten policy families from five vendors (30,942 episodes with a simulated customer), when the false value licenses a refusal, a correction one turn in buys back 0.678 [0.602, 0.754] of the task success the false premise cost and four turns in only 0.254 [0.132, 0.376] (pooled over the nine families the fault damages, on the same units). This decline, the correction window (success corrected at minus ), is positive in all six API-served families, where it pools to +0.287 [+0.217, +0.358] and to +0.294 [+0.210, +0.377] net of the same edit applied to a context with nothing false in it; net of that edit, none is detected in the four open-weight families we served locally. Withholding the same record until the same turn shows that lateness is part of the window but not all of it. Among episodes that reach the late correction the window is smaller under a second simulated customer, though not when every scheduled episode is counted. Late resets of the agent's side do not reliably restore success, and an explicit retraction adds no reliable benefit over silent correction. All code, data and episode traces are in the supplementary material and will be released at https://github.com/xxx/xxx upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.