acceptodds
Under review as a conference paper at ICLR 2027

ToolTailor: Personalised Tool Use in Multi-Turn Travel Planning

Abstract

Tool-augmented language models have become the dominant paradigm for task-oriented assistants, yet existing benchmarks rarely evaluate whether these models integrate individual user preferences into their tool calls. We introduce ToolTailor, a benchmark for personalised tool use in travel planning, together with a synthetic data generation pipeline that produces personalised tool-use trajectories at scale. ToolTailor is the first benchmark to combine long-horizon, multi-turn tool use with rich user profiles, requiring agents to satisfy individual preferences while orchestrating multiple tools across a task. Tasks are seeded from travel-planning scenarios collected from web forums, and the development and test splits are human-validated. Evaluation with state-of-the-art LLMs shows that personalised tool use remains challenging: models fail to identify and incorporate relevant user preferences into tool calling. Supervised fine-tuning on our synthetic trajectories improves preference integration by up to 35 F1 points when the relevant preferences are provided directly, while also improving schema compliance. When the model must instead identify them in the full user profile, all four fine-tuned open models still outperform three proprietary models (Gemini-3.6-Flash, GPT-4o-mini and GPT-5.6-Terra) prompted zero-shot.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.