Anatomy of a Split: What One Worker Per Function Costs
Abstract
Orchestrator-worker systems give each tool family its own LLM specialist. We ask, under preregistration, what that split costs in function calling and whether the loss can be repaired. The split we study is the simplest: a planner assigns each function a request needs to a same-model worker, each worker emits that function’s calls, and the lists are concatenated, under a rule we set ourselves that keeps only the first item when the planner names the same function more than once; the comparison is one monolithic call (BFCL v4; Gemma-3-12B, Gemma-3-27B, Llama-3.3-70B; a sealed test set). Under the registered wording no split configuration we ran is shown to beat one call; under a rewording that costs the monolithic prompt 11.3 points at 12B, SPLIT+PASS sits 11.9 above it. The registered split loses 13.4, 10.8 and 3.9 points. A control that undoes our first-item rule shows that rule to be the largest single class of loss on the Gemma models at point estimate, and inert on Llama; the registered test of whether the rule is the majority of the harm was inconclusive on the Gemma models. The remaining loss is concentrated where the planner actually splits the request (a post-hoc cut). The one repair confirmed under preregistration, a further call that sees the whole request and the split’s calls and corrects them, is confirmed on Gemma-3-27B only; its Gemma-3-12B reading met the registered criterion but hinges on single tasks, so we do not claim it, and on Llama it was not met. It is not shown to need the split: the split’s calls add at most two points to it on the Gemma models by unadjusted intervals, and at 12B the same call given nothing to correct is better by an interval wholly below zero (inconclusive on Llama). No repair that keeps the split is confirmed. Post hoc, the loss is planner items our rule drops, wrong function sets and argument errors, and at a large catalogue the loss grows, mostly because a planner shown only names and descriptions picks near-duplicate tools. Except where marked, the rest is descriptive or post hoc; the control-adjusted loss at 27B does not replicate on our earlier test split; the four arms the claims rest on ran under three prompt wordings on all three models, where the control-adjusted loss keeps its sign except at 12B, whose monolithic prompt loses 11.3 points to one rewording. Scope: one-shot text emission, one benchmark family, 1 to 4 functions per task; our registered run inside a native tool-calling API came out refuted and was discarded afterwards for corrupted payloads.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.