ToolNeed: Selective Verification with Matched No-Tool Behavior
Abstract
Tool augmentation changes language models from generators into systems capable of initiating external actions. Before tool selection or argument generation comes a more fundamental decision: is external action warranted at all? We distinguish tool validity from behavioral necessity: a proposed action may be relevant, valid, executable, and even lead to a correct outcome, yet remain unnecessary relative to what the same model can accomplish without tools. In targeted construct analyses, matched no-tool behavior remains correct for 96.3% and 99.4% of successful valid tool paths on Qwen3.5 and DeepSeek, respectively. We introduce ToolNeed, a training-free, black-box-compatible method that selectively verifies proposed tool actions by eliciting matched no-tool behavior from the same target model and using it for contrastive arbitration. Candidate-blind routing determines when this verification is activated without conditioning on the concrete proposal. On SABEval, aggregate final-action Tool Invocation Rate falls from 91.2% to 8.1%, 43.0% to 8.5%, and 62.1% to 18.0% across three backbones. On the shared BFCL held-out Live subset and Non-Live split, ToolNeed achieves the highest Irrelevance accuracy in all six backbone–regime settings. Tool-required behavior is near-exactly preserved on Non-Live, while Live settings show modest preservation losses.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.