NavMind: From Open-Ended Navigation Requests to Grounded Personalized Routes
Abstract
Conventional route planners optimize distance or time but rarely jointly consider visit needs and preferred street-level experience. Planning with large language models (LLMs) alone seems strong in open-ended requests, but may hallucinate places, misjudge street environments, or violate road-network constraints. We present NavMind, an LLM-guided personalized routing framework that grounds visit intent (e.g., a photogenic restaurant serving good steak) and preferred route experience (e.g., lively, visually engaging streets) in separate evidence sources using three modules. (1) Hybrid Point-of-Interest Retrieval-Augmented Generation (Hybrid POI-RAG) encodes POI locations and review-derived semantics to identify POIs that match the visit intent. (2) The perceptual-cost module scores street qualities with a fine-tuned vision-language model, while LLM-as-a-Jury reconciles different LLM estimates of each quality’s importance to the requested experience. (3) The route-optimization module casts visit intent as waypoint constraints and preferred street-level experience as edge costs in a generalized traveling salesman problem, yielding a network-valid route. Evaluation on Beijing data spans diverse personas and requests. Qualitative cases illustrate how NavMind adapts waypoints and routes to traveler preferences, generating diverse routes across intentions. At scale, simulations indicate that this adaptation requires only modest detours, with median detour rates below 20 for every persona. In user assessment, the resulting routes score 1.78 points higher than shortest paths on a seven-point scale and are preferred nearly four times as often. Component-level results further support these system-level findings: Hybrid POI-RAG achieves over five times the Precision@5 of semantic-only or spatial-only baselines; street-view scores agree with human ratings across four qualities (Pearson’s > 0.90); and Jury profiles align more closely with human judgments than single LLM. These results highlight NavMind’s potential for meaningful personalized navigation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.