acceptodds
Under review as a conference paper at ICLR 2027

When Does Model-Specific Tool-Call Formatting Pay Off?

Abstract

Different LLMs often prefer different tool-call formats. Should each model then get its own format? Our framework, FormatGain, separates the model–format in- teraction, the oracle gain of per-model formats over the best shared format, and the selection error of a small-budget rule. Decision margins link the first two: a model gains from another format only when its interaction advantage there ex- ceeds that format’s average deficit. So one interaction can come with a large gain or none, and a gain can still be lost in selection. Eight designs, five on our suite and three external, show three regimes (post hoc): little to gain, a gain lost, a gain kept. On WorkBench, the corrected interaction share is 0.38 but the oracle gain is only 0.001, and fixed json and a rule pooling the other models’ history both had lower observed regret than every probe. On three related designs of our suite, a ten-task probe gains 0.030 to 0.105 over the hindsight best shared format. The gains transfer only in part: with new models they held on the strict-parser and 4k-token designs but not on three others, they held on new task templates, and on new BFCL task families the probe showed no detected advantage over the pooled-history rule.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.