acceptodds
Under review as a conference paper at ICLR 2027

CAFTBench: A Benchmark for Continuous Fine-Tuning of Tool-Calling Agents

Abstract

LLM-based agents increasingly rely on external tools and APIs, yet most deployed systems assume a static tool ecosystem. In practice, APIs evolve: new tools appear, schemas change, and existing parameters are modified. Such changes require agents not only to acquire new tool-use behavior, but also to retain earlier tools, adapt to schema mutations, and compose skills learned at different times. We introduce **CAFTBench**, a static and reproducible benchmark for evaluating continuous fine-tuning of tool-calling agents. CAFTBench consists of **CAFT-30k**, a tool-ordered training corpus of 29,573 tool-calling instances spanning more than 260 tools, and CAFT-Eval, a probe suite covering single-tool routing, cross-stage composition, API upgrades, no-tool abstention, and structured-output robustness. We further define five metrics—Historical Routing Failure (HRF), Cross-Stage Composition Rate (CSCR), Negative Transfer and Historical Interference (NTHI), Tool Over-reliance Index (TOI), and Syntax Collapse Rate (SCR)—to disentangle retention, transfer, interference, over-reliance, and output-format degradation. Experiments with Qwen3-8B and LLaMA-3.1-8B-Instruct show that existing continual learning methods exhibit strong backbone-dependent trade-offs: EWC-style regularization improves stability and upgraded-API adaptation on Qwen, while semantic replay improves composition, but the same methods can substantially degrade LLaMA and induce tool over-reliance. These results highlight the need for continual learning methods designed specifically for evolving tool-calling agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.