ToolTailor: Personalised Tool Use in Multi-Turn Travel Planning
Abstract
Tool-augmented language models have become the dominant paradigm for task-oriented assistants, yet existing benchmarks rarely evaluate whether these models integrate individual user preferences into their tool calls. We introduce ToolTailor, a benchmark for personalised tool use in travel planning, together with a synthetic data generation pipeline that produces personalised tool-use trajectories at scale. ToolTailor is the first benchmark to combine long-horizon, multi-turn tool use with rich user profiles, requiring agents to satisfy individual preferences while orchestrating multiple tools across a task. Tasks are seeded from travel-planning scenarios collected from web forums, and the development and test splits are human-validated. Evaluation with state-of-the-art LLMs shows that personalised tool use remains challenging: models fail to identify and incorporate relevant user preferences into tool calling. Supervised fine-tuning on our synthetic trajectories improves preference integration by up to 35 F1 points when the relevant preferences are provided directly, while also improving schema compliance. When the model must instead identify them in the full user profile, all four fine-tuned open models still outperform three proprietary models (Gemini-3.6-Flash, GPT-4o-mini and GPT-5.6-Terra) prompted zero-shot.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.