ToolForm: A Controlled Study of Tool Formatting for Agents
Abstract
Agent frameworks show a language model its tools in one written format, such as JSON Schema, XML, YAML, a Python function signature or a natural-language description. Prior studies compare these formats as one design choice and reach different conclusions. A tool interface sets three things: the schema format the model reads, the answer protocol it writes back in, and the grading rule that scores the call. Most prior studies change the first two together and fix the third, so the cause of a difference is unclear. We introduce **ToolForm**, an open-source framework that separates the three into independently controlled variables, measures the effect of each, and turns the measurements into defaults an agent can use. Our results show that many apparent format advantages depend on the answer protocol and the grading rule, while the gain from showing the parameter schema keeps its direction under every rule and task set. **First**, we run a controlled study that changes one factor at a time, on **100 open-source models** in 17 families (0.36B to 122B) and **12 proprietary models**, with **129 schema formats**, **31 answer protocols**, **3 grading rules**, 1 to 256 tools, 5 complexity tiers, 6 task regimes and 3 pressure treatments, at 5 seeds. **Second**, we build a benchmark from three sources: a seeded generator over **24 domains** whose tools run against a world state, **66 real tool schemas** with **1,248 human-written requests**, and **6 public benchmarks** rewritten in our formats and scored by their own checkers. An audit checks every record, and two authors check **600 graded turns** by hand. **Third**, we package ToolForm as a toolkit whose advisor turns the results into a format and protocol choice for a given model and tool set. The results give four rules. *Show the parameter schema*: it raises accuracy by **57.09%** over tool names alone, the largest effect in the study. *Choose the answer protocol before the schema format*: the protocol changes accuracy **1.22 to 2.00 times** as much as the format, and the XML lead over JSON Schema under the default pairing (13.80%) falls to 2.88% over five protocols. *Prefer a structured format to natural language*: natural language is never more accurate on average than the best structured format at any model size. *Keep the tool list short and treat tool descriptions as untrusted*: going from 1 to 256 tools lowers accuracy by 41.01%, and an instruction hidden in a tool description is followed on **38.48 to 54.16%** of turns in every format that shows the description. The gap between formats narrows with model size but does not close: the median of the 20 anchor models (12 proprietary) varies by 25% across the eight core formats, against 44% for the median open-source model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.