acceptodds
Under review as a conference paper at ICLR 2027

Calling Is Not Using: Do Post-Training Gains Increase Dependence on Tool Returns?

Abstract

Tool-using language models are judged by task success, which does not separate using a tool from using what the tool returns: a model can score higher by calling more often, executing more reliably, and stopping in a scorable format, none of which requires reading a return. We trace interactions through five stages: routing, execution, acquisition, integration, and termination. Only integration leaves no trace, since a response can agree with a return without being caused by it; we measure it by intervention instead. Replaying each episode twice from the same history and first action, Natural replay preserves tool returns while matched-null replay replaces every executed return with a length-matched uninformative observation. Their paired accuracy difference measures the return’s contribution to task success, and its change relative to the frozen base model is what post-training did to that dependence. Across our training grid, returns are worth most on the interface a model trained against: elsewhere accuracy rises, but most of the gain survives deleting them. Transfer does occur: among the single-source settings, supervised fine-tuning on mathematics yields the only confirmed increase on a debugging interface where every policy starts from the same failing program. But acquiring a return is not using it: on a diagnostic answerable only from the retrieved passage, all three Search-SFT runs retrieve the answer in every example yet answer only 7.3% of items correctly; given it in context, they are almost always right. Scoring well, acquiring information, and using it are three different things; tool-use transfer should be tested by the contribution of returns to success, not inferred from task performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.