acceptodds
Under review as a conference paper at ICLR 2027

Interpreting and Improving Supervised Fine-Tuning for Tool Use in LLM Agents via Interactions

Abstract

Supervised fine-tuning (SFT) improves the tool-use capabilities of large language model (LLM) agents, yet it remains unclear why tool-use capability varies with the SFT-based method, the model scale, and the amount of SFT data. With Harsanyi interactions, we investigate whether a common mechanism underlies these factors: for the two stages of tool calling, tool selection and argument generation, we decompose the LLM's output score into interaction effects among input words, where complex interactions involve many words and simple interactions involve only a few. We find that SFT, more effective SFT-based methods, larger model scales, and more SFT data all tend to strengthen the LLM's modeling of complex interactions and weaken its modeling of simple interactions. Motivated by this common mechanism, we propose Interaction Complexity Regularization (ICR), a plug-and-play auxiliary loss that directly enhances the strength of complex interactions while reducing that of simple interactions during SFT and can be added to existing SFT-based tool-use methods for tool selection and argument generation. Experiments with various backbones from five LLM families on three tool-use benchmarks, BFCL, StableToolBench, and -bench, show that ICR enhances SFT across model families and SFT-based tool-use methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.