acceptodds
Under review as a conference paper at ICLR 2027

Lightweight Tool Probing: Showing LLM Agents What Their Tools Return

Abstract

Tool documentation tells an agent how to call a tool, but only calling it shows what it returns. Multi-turn function calling runs on exactly this missing knowledge. Arguments must match the formats a tool expects, and identifiers returned by one call must feed the next. Existing remedies either keep the agent reading, rewriting the documentation into better text about the tool, or keep it paying, rediscovering tool behavior inside every task; and task success alone cannot show whether such guidance helps. We introduce Lightweight Tool Probing (LTP): before any task, a model probes a toolset once in an isolated sandbox and keeps the raw responses it observed in a frozen evidence bank; during a task, it retrieves at most three short response excerpts per step, while identifiers still come from the live environment. LTP needs no gold solutions, no task feedback, and no parameter updates. To locate where the evidence helps, we also introduce execution diagnostics that score how much of a task’s requirements an execution covers. On BFCL, LTP lifts DeepSeek-V4-Flash from 41.75% to 49.75% at 1.09 times the tokens of plain tool calling, outperforming all five baselines, the strongest of which need at least twice the tokens. The gain travels to GLM-5.3 and Qwen3.5-397B, and to ACEBench, AppWorld, and WorkBench. It is interpretable. The tasks LTP rescues gain argument-constraint coverage. Even its trajectories are better training data. They beat plain trajectories for all four student models, by up to 10.75 points. **Showing an agent what its tools actually return is a cheap complement to telling it how to call them.**

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.