acceptodds
Under review as a conference paper at ICLR 2027

TA-CallBench: Benchmarking Nested Tool Calling for Conditional Trigger–Action Tasks

Abstract

Large language models (LLMs) can select tools and generate arguments from user requests, supporting immediate queries, software operations, and task automation. However, existing benchmarks primarily evaluate whether models generate correct function calls for the current request, and rarely examine their ability to generate nested tool calls for conditional trigger–action tasks, where an operation must be executed only after a specified event or state occurs. Because such a task combines triggering conditions with executable Actions, its representation requires an outer call that organizes these components. To evaluate this ability, we introduce TA-CallBench, which requires a model to select inner tools (the individual Trigger and Action functions) and outer tools (management tools that organize inner calls into complete conditional tasks) from the available tool set and generate the required nested tool calls. TA-CallBench provides thousands of bilingual dialogues across six application scenarios and 12 task types, including single-task creation, multi-task creation, multi-turn modification, and fine-grained component editing, which together require a model to identify task boundaries and organize components within nested calls. We construct the data through Trigger-anchored tool-combination sampling, LLM-based generation, structural validation, and combined model and human review. The evaluation centers on Tool-Call Accuracy (TCA), which checks whether a model correctly generates the Trigger and Action calls and their arguments. All evaluated models achieve TCA below 55% on TA-CallBench. For models evaluated on both benchmarks, their TA-CallBench TCA is substantially lower than their performance on the Berkeley Function Calling Leaderboard (BFCL). This gap indicates that current LLMs still struggle to reliably generate such nested calls, highlighting the need for dedicated evaluation of this capability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.