ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation
Abstract
Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior. Existing function-calling benchmarks often emphasize successful, cooperative interactions and underrepresent adversarial conversation trajectories, thereby limiting the training resources available for developing robust agents. We present ToolRACER, a synthetic data generation pipeline that coordinates a user, assistant and tool agent to generate and validated multi-turn interaction between user and agent. Using ToolRACER, we construct ToolRACER-Bench a robust multi-turn conversation benchmark spanning across six domains, ranging over 55 varied personas, generating a validated corpus of 5.8K conversation trajectories, with approximately 66% of conversations containing failure-prone conversation scenarios. We inject adversarial trajectory behavior, producing validated conversation interaction trajectories that capture realistic, robust scenarios. We evaluate models trained on ToolRACER-Bench against internal benchmark as well as function calling benchmarks such as -bench and ACEBench to evaluate agentic accuracy and agentic robustness. Models trained on ToolRACER-Bench improves end to end agentic accuracy across -bench, BFCLv3 and ACEBench, demonstrating significant gains when mixed with in-domain dataset in small language models for agent capability tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.