acceptodds
Under review as a conference paper at ICLR 2027

Correct Labels Are Not Enough: Which Part of a Judgment Keeps a Tool Agent Safe?

Abstract

Agents often pass on only the verdict of an earlier check. We ask whether correct labels kept alone still stop an agent from acting on a false claim in its context. In the airline task of τ -bench, a colleague’s note falsely says that the airline cancelled a flight, which would allow a refund; we keep parts of the agent’s saved assess- ment and grade the final database. When gpt-oss-120b judged the note FALSE and planned to KEEP the booking, its whole assessment led to no forbidden cancella- tion and its two labels alone to 76%; all ten agents of our panel lose protection. Our thesis is that agents act on stated facts more than on verdicts, and weigh an- other agent’s message by who is said to send it and whether it says it was checked. One sentence of the fact restores most of the protection wherever it stands, a con- clusion does not, and the loss also appears when the claim favours refusal. The same stop instruction protects more from another agent than from the agent itself in all five agents we test, and a verdict protects more when it claims a check. Our evidence comes mainly from one airline task, where the record only implies the fact; the two strongest models we tried ignore the note. Systems should check permission against the database and pass a checked premise, not only a verdict.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.