Remembering Is Not Navigating: Goal Recurrence and Closed-Loop Success in City-Scale Drone Navigation
Abstract
Vision-language navigation agents are increasingly fine-tuned on the benchmarks that rank them, and a higher closed-loop success rate is read as evidence that an agent has learned to navigate. Success can also come from memory: if the places an agent is sent to also appear in its training data, recalling where a description led is enough. EmbodiedNav-Bench, a city-scale drone benchmark, is exposed to this risk: it ships without a train/test split, so each fine-tuning study draws its own, and nothing checks what the held-out goals share with training. We measure how much success memory alone can buy there. Holding out every tenth episode, we find that held-out goals are rarely new: for 60 of them, one training episode has both a near-duplicate instruction and a goal inside the success radius. On 101 held-out episodes flown in the benchmark's simulator under our harness, a controller that flies from the start pose to the remembered endpoint of the matching training instruction succeeds on 83 of the episodes it flies from memory and on 58 overall, with a learned policy acting on the rest, against 22 for an imitation-learned policy and 32 for always flying forward; without a near-duplicate instruction or a nearby training goal, it does no better than flying forward. Gains from training fixes need the same scrutiny: an offline exposure-bias correction with geometry-verified preference optimization reduces the harm of the agent's own action history in offline next-action accuracy, yet same-budget controls without the preference signal give reductions we cannot distinguish from it, and its closed-loop effect is untested. We release goal-disjoint episode lists and recommend reporting success on them beside a retrieval baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.