When Are Tool Calls Necessary? Counterfactual Evaluation and Pre-Generation Epistemic Gating in Tool-Augmented Language Models
Abstract
When a tool-augmented language model outperforms its closed-book counterpart, the improvement is routinely credited to external computation or retrieval. We show that standard evaluations conflate two effects: the information a tool provides (Non-Parametric Information Gain, NPIG) and the effect of switching to a protocol in which the model must format a tool call before answering (Structural Scaffolding Effect, SSE). To separate them, we introduce an 8-Arm Counterfactual Matrix that holds Turn-1 tool calls byte-identical across real execution, self-simulation, schema-echo, empty return, and adversarial conflict. On solvable parametric tasks, action-first JSON calling lowers accuracy below closed-book answering. The largest group of lost items has a valid call but no checking after the tool returns (Stage-2 Post-Call Deliberation Suppression); smaller groups have call arguments corrupted before reasoning (Stage-1 Premature Schema Compression) or no call at all. Control experiments show that the loss is not caused by showing tool schemas or by our echo wording: it appears once a call is required, persists with neutral echo text, native tool templates, and grammar-constrained JSON, and mostly disappears when the model is asked to verify its answer after the tool returns. In a 40-item-per-benchmark study across 12 SOTA datasets with a local lookup/solver tool, an oracle that returns the gold answer on every call gives 29.9 percentage points higher accuracy than real execution; on the calls that return the gold answer, text LLMs jump by 55.7 percentage points over an echo (), whereas Qwen2.5-VL () already reads the answer from the image closed-book and barely changes across real, echo, or empty returns. Finally, our truncation-corrected entropy () predicts closed-book errors on lookup tasks, but requires per-task calibration and does not beat simpler prefix scores on actual tool benefit; when tools are fallible, prefix uncertainty is near chance and our passive 90-feature Logit-Lens probes fall to chance on unseen models and benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.