acceptodds
Under review as a conference paper at ICLR 2027

TriRoute: Routing Tool-Call Failures Across Resampling, Stopping, and Editing

Abstract

The structure of a failed tool call can reverse the value of its next recovery action: diffuse malformed outputs favor fresh samples, concentrated syntax-valid faults favor local edits, and low-separation states can make stopping preferable. TriRoute operationalizes this geometry with a frozen six-feature score that selects among Best-of-, Keep, and a radius-three TriEdit search. To distinguish the value of the resulting policy from a stronger repair arm, routed and greedy policies share the generator, verifier, prompts, seeds, five-template edit library, and a hard ceiling of six verifier evaluations. With GPT-4o, TriRoute improves over always invoking that identical edit library by – percentage points on APIBench-Tool, BFCL v3, and API-Bank, with paired intervals excluding zero, while using – fewer verifier evaluations per recovery episode. Against schema-aware validator-feedback retry, it reaches on BFCL v3 and on API-Bank—gains of and points—and reduces recovery-episode calls from to . Score strata and failure slices exhibit the predicted shift from resampling to editing; without generator-specific fitting, routed-minus-validator intervals exclude zero in all nine cells spanning GPT-4o, Llama-3.1-70B, and DeepSeek-32B, while same-library point estimates are positive in all nine and six intervals exclude zero. In verifier-rich structured generation, assigning an existing recovery repertoire from the observed failure state raises success while reducing verification demand, without enlarging the repair library.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.