OUTPUT FORMAT AND INDEX LOOKUP DEGRADE INCONTEXT TOOL SELECTION AT LARGE CATALOG SIZES
Abstract
When an LLM selects a tool from a large in-context catalog, how much of its degradation with catalog size is due to output format and index lookup, as opposed to selection itself? Across Qwen2.5 (0.5B–14B, 4-bit) on the Berkeley FunctionCalling Leaderboard (BFCL), a zero-learning lexical-overlap heuristic overtakes in-context LLM selection at a finite, measurable catalog size |V | ∗ . This crossover grows with scale (8.7 at 1.5B to 500.0 at 14B; α = 1.784), then more slowly beyond 14B (365.3 at 32B, upper bound resolved rather than right-censored; 552.6 at 72B, 95% CI [346.6, >650], upper bound right-censored — both precisionmatched, App. B.24). Two frontier closed models (GPT-4o, Claude Sonnet 5, small-n) show no crossover at all through |V |=650 (App. B.26, B.27), limiting practical generality. Forcing name-based rather than index-based output recovers a substantial share of the lost performance at high |V | (3B: +37.5pp, 7B: +41.5pp, both at |V |=234; 3B peaks at +50.5pp at |V |=64), but performance still declines with |V | under name emission, so selection difficulty and output-format cost are not cleanly separated — a direct log-likelihood control that removes output format entirely does not reveal an easier underlying selection ability, if anything the reverse at lowto-moderate |V | (App. B.21). Extending the grid to |V |=650 resolves 3B’s own name-emission crossover (|V | ∗=350, 12.2× index-emission’s point estimate, CI [9.4×, 15.1×]); 7B stays right-censored, so its ratio is an upper bound only, since a hashed-name control could not cleanly rule out lexical-shortcut copying as the source of the gain. A further test removes the task entirely, asking only whether the model can locate a named, given answer’s index: accuracy still declines significantly with |V | at both scales (3B: −70.0pp; 7B: −41.5pp), though at 3B the task is counterintuitively harder than ordinary selection at |V |≤16 — near-miss counting fits well (74.2%, 2.6× chance, at 8; 41.0%, 3.1× chance, at 16), not a bug. Two further controls narrow the picture: |V |, not raw prompt length, is the primary driver (partially so at 3B), and the near-miss pattern is not specific to one binding direction; we do not claim a fully clean causal isolation of this mechanism. A second mitigation, filtering the catalog before selection, closes most of the gap, and is effective across a range of k rather than only at a model’s own crossover (App. B.9). Most of this gain is achievable through candidate ranking alone, without truncation — a cheaper, more directly deployable finding, interacting with the output-format effect (§4.5). Every |V | ∗ reported here is specific to lexicallytransparent catalogs: a direct opacity check collapses the crossover roughly 3× (App. B.7), and the overall finding replicates directionally on an independent benchmark and model family. We report every number here, including the ones that complicate the story, as measured.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.