Beyond Functional Correctness: Optimizing Multi-turn Tool-Augmented Agent Trajectories to Reduce Tool Usage
Abstract
Large language model (LLM) agents are increasingly evaluated on multi-turn tool use, where success depends not only on calling the correct APIs but also on keeping interaction trajectories compact. Existing alignment methods largely optimize for correctness, while efficiency-oriented approaches mainly target retrieval or single-query settings, leaving open how to reduce excess tool calls in sustained, stateful dialogues without sacrificing task success. We propose an iterative Direct Preference Optimization (DPO) framework that improves both accuracy and efficiency for multi-turn tool agents. For each query, we sample multiple successful trajectories, rank them by tool-call overhead, and convert best–worst discrepancies into step-level preference pairs via two complementary rules: one detecting redundant or corrective detours, and another recovering divergences under parallel tool calling. A greedy decoding filter retains only pairs the current policy has not yet mastered, and the sample–construct–optimize loop is repeated to expose deeper inefficiencies inside turns. When this cross-trajectory loop plateaus, we further mine preferences from the greedy successful trajectory via validated pruning—ablating removable tool calls under the BFCL multi-turn checker—and apply an additional DPO stage with cleaned prefixes. On BFCLv3 multi-turn subsets, LoRA-based DPO followed by this prune stage raises an 8B BitAgent model to 79.5% average accuracy, surpassing the previous overall state of the art (xLAM-2-70b-fc-r, 77.38%), while reducing excess tool calls on the Base subset from 2082 to 25. Trained only on Base, the same policy generalizes to Miss-Param, Missing-Func, and Long-Context subsets with fewer excess tool calls (and often higher accuracy), and also lowers total tokens and inference steps, indicating that compact trajectories often coincide with correct multi-turn behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.