acceptodds
Under review as a conference paper at ICLR 2027

Please Put Nothing Here: When Tool Documentation Helps LLMs Make the Wrong Call

Abstract

In tool interfaces, leaving a field out, sending null, sending an empty value, and sending the default can mean different things, depending on each tool's documented rules. A model can therefore send a call that is valid and executes, yet writes a state the user did not ask for. We introduce \bench, which pairs synthetic partial-update tasks that differ in one documented rule and scores the exact resulting state by deterministic execution. In aligned tasks the most direct call is correct. In conflict tasks the rule makes that call write a predictable wrong state. We test whether documenting every input form improves accuracy over stating only the rule the task depends on. Across 144 held-out task families (each an aligned and a conflict version of one task) and five model configurations, a table of equal length that gives the operation for all ten ways of writing a field raises aligned accuracy over a one-line rule but lowers conflict accuracy from 82.2% to 76.7%, below even a schema that gives no rule for omitted, null, empty, or default inputs (82.9%). The difference-in-differences (the table's effect on conflict tasks minus its effect on aligned tasks) is points and negative for all five models. Accepted calls that reach the predicted wrong state rise from 19 to 77 of 720 conflict responses, while invalid calls fall from 44 to 18 and aggregate accuracy barely moves. Four-row tables recover much of the loss. On four newer models, the table's effect is smaller on average and varies by model. In a small case series on three protocol rules that depart from JSON Merge Patch, a short table scores higher than the one-line rule. Documentation changes should therefore be evaluated on paired conflict tasks and on the states they produce, not only on call validity or aggregate accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.