acceptodds
Under review as a conference paper at ICLR 2027

NavGround3D: UAV-Centric Grounding of Semantic–Geometric Cues for Aerial Vision-Language Navigation

Abstract

Aerial Vision-and-Language Navigation (UAV-VLN) requires UAVs to follow natural-language instructions, understand scene semantics and spatial relation- ships in complex 3D environments, and navigate autonomously. Recent studies have begun to leverage geometric visual models in UAV-VLN to enhance 3D spa- tial understanding. However, existing approaches mainly focus on aligning or fus- ing semantic and geometric information, while paying limited attention to effec- tive visual information distillation and unified UAV-centric directional modeling. To address these issues, we propose NavGround3D, a semantic-geometric rep- resentation framework that progressively distills visual information and grounds it in a unified UAV-centric directional space. Specifically, Semantic-Geometric Visual Consolidation (SGVC) selects representative anchors based on seman- tic saliency and aggregates redundant representations according to semantic- geometric consistency, yielding more informative visual tokens. UAV-Centric Di- rectional Grounding (UCDG) then transforms view-dependent visual cues into the UAV body frame and introduces continuous directional encoding to estab- lish consistent cross-view spatial relationships. The resulting representations are integrated with navigation context for motion decision-making. Extensive exper- iments on TravelUAV and UAV-ON demonstrate consistent improvements over strong baselines, validating the effectiveness of semantic-geometric distillation and UAV-centric directional grounding for aerial navigation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.