acceptodds
Under review as a conference paper at ICLR 2027

OUTPUT FORMAT AND INDEX LOOKUP DEGRADE INCONTEXT TOOL SELECTION AT LARGE CATALOG SIZES

Abstract

When an LLM selects a tool from a large in-context catalog, how much of its degradation with catalog size is due to output format and index lookup, as opposed to selection itself? Across Qwen2.5 (0.5B–14B, 4-bit) on the Berkeley FunctionCalling Leaderboard (BFCL), a zero-learning lexical-overlap heuristic overtakes in-context LLM selection at a finite, measurable catalog size |V | ∗ . This crossover grows with scale (8.7 at 1.5B to 500.0 at 14B; α = 1.784), then more slowly beyond 14B (365.3 at 32B, upper bound resolved rather than right-censored; 552.6 at 72B, 95% CI [346.6, >650], upper bound right-censored — both precisionmatched, App. B.24). Two frontier closed models (GPT-4o, Claude Sonnet 5, small-n) show no crossover at all through |V |=650 (App. B.26, B.27), limiting practical generality. Forcing name-based rather than index-based output recovers a substantial share of the lost performance at high |V | (3B: +37.5pp, 7B: +41.5pp, both at |V |=234; 3B peaks at +50.5pp at |V |=64), but performance still declines with |V | under name emission, so selection difficulty and output-format cost are not cleanly separated — a direct log-likelihood control that removes output format entirely does not reveal an easier underlying selection ability, if anything the reverse at lowto-moderate |V | (App. B.21). Extending the grid to |V |=650 resolves 3B’s own name-emission crossover (|V | ∗=350, 12.2× index-emission’s point estimate, CI [9.4×, 15.1×]); 7B stays right-censored, so its ratio is an upper bound only, since a hashed-name control could not cleanly rule out lexical-shortcut copying as the source of the gain. A further test removes the task entirely, asking only whether the model can locate a named, given answer’s index: accuracy still declines significantly with |V | at both scales (3B: −70.0pp; 7B: −41.5pp), though at 3B the task is counterintuitively harder than ordinary selection at |V |≤16 — near-miss counting fits well (74.2%, 2.6× chance, at 8; 41.0%, 3.1× chance, at 16), not a bug. Two further controls narrow the picture: |V |, not raw prompt length, is the primary driver (partially so at 3B), and the near-miss pattern is not specific to one binding direction; we do not claim a fully clean causal isolation of this mechanism. A second mitigation, filtering the catalog before selection, closes most of the gap, and is effective across a range of k rather than only at a model’s own crossover (App. B.9). Most of this gain is achievable through candidate ranking alone, without truncation — a cheaper, more directly deployable finding, interacting with the output-format effect (§4.5). Every |V | ∗ reported here is specific to lexicallytransparent catalogs: a direct opacity check collapses the crossover roughly 3× (App. B.7), and the overall finding replicates directionally on an independent benchmark and model family. We report every number here, including the ones that complicate the story, as measured.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.