AerialGenesis: Language-Conditioned Predictive Safety Fields for Generative UAV Navigation
Abstract
Language-guided navigation of aerial robots requires an agent to interpret open-vocabulary instructions, anticipate partially observed consequences, and remain safe around unknown semantic hazards. Existing UAV reinforcement-learning systems are predominantly reactive and geometric, while vision-language navigation and vision-language-action models rarely provide a deployable safety mechanism. We introduce AerialGenesis, a unified formulation of language-conditioned safe generative navigation. A predictive semantic world model encodes observation histories and instructions into a latent state that forecasts visual features, semantic occupancy, navigation affordances, collision probability, and task completion. A risk-aware flow-matching policy then generates trajectory distributions rather than a single action, explicitly trading task alignment against predicted risk and uncertainty. A self-evolving neural safety field is learned from demonstrations, failures, and interaction data; it replaces brittle hand-designed barrier functions while retaining a probabilistic certificate under calibrated uncertainty. We derive a bound connecting world-model error, safety-field calibration, and the collision probability of sampled trajectories. We evaluate the method on simulated and real-world UAV navigation benchmarks covering language-guided navigation, exploration, semantic search, unknown obstacles, and zero-shot sim-to-real transfer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.