acceptodds
Under review as a conference paper at ICLR 2027

MobiLens: A Human-Centered Benchmark for Personalized Mobility Agents

Abstract

Large language models have enabled mobility agents to interpret complex travel intents, use navigation tools, and tailor route and destination recommendations to individual users' needs, preferences, and contexts. However, existing benchmarks provide limited support for jointly assessing behavioral alignment and broader human-centered outcomes under a common evaluation protocol. To address this gap, we introduce MobiLens, a human-centered benchmark built from real-world user queries, pre-query user context, and observed mobility decisions. Inspired by naturalistic decision-making theory, our framework evaluates mobility agents across five human-centered outcome dimensions: personalization, diversity, fairness, explainability, and robustness, alongside process-level measures of tool-use success and operational efficiency. We benchmark 11 agent configurations spanning reasoning, memory augmentation, multi-agent coordination, and recommendation-oriented designs under a shared offline navigation-tool replay protocol. Our results reveal an execution–alignment gap: most agents reliably execute tools and produce format-compliant answers, yet show limited alignment with users' observed route and destination choices. Performance also varies across user groups and mobility scenarios, underscoring the need to evaluate behavioral alignment beyond aggregate execution success. Code and benchmark resources are available at https://anonymous.4open.science/r/MobiLens-F917.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.