LODGE: A CONVERSATIONAL TRAVEL AGENT BENCHMARK WITH GROUNDED GOALS AND CONTROLLABLE DIFFICULTY
Abstract
Benchmarks for conversational search agents typically assume a highly specific target: the user knows exactly which flight to buy or which hotel to book. Real users usually arrive with underspecified requests. Evaluating such agents therefore requires tasks whose correct outcome is known, with difficulty that can be varied deliberately rather than incidentally. We present LODGE, a benchmark of 360 conversational accommodation-booking tasks, 300 of them a fixed evaluation split on which all results here are reported, each with a unique bestfitting target listing established by contrastive construction and confirmed by human review, over a shared catalog of synthetic listings. Difficulty is expressed as composable axes that vary one property of the task while the definition of success stays fixed. GEOFLEX varies whether the named destination is the target locality (PRECISE), an anchor to search beyond (NEARBY), or a region to refine into localities (BROAD); VAGUENESS varies how explicitly the same goal is disclosed. Across seven language agents, these axes separate capabilities that aggregate accuracy conflates: booking completion is nearly uninformative about task success, NEARBY is harder than BROAD for every agent despite supplying more locational information, and robustness to vagueness is independent of geographic competence. Small-scale supervised fine-tuning improves the difficult settings disproportionately, with results consistent with transfer to held-out countries. The dataset, user simulator, and evaluation harness will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.