acceptodds
Under review as a conference paper at ICLR 2027

Spatio-Temporally Consistent Panoramic Generation for Vision-Language Navigation

Abstract

The generalization of Vision-and-Language Navigation (VLN) agents remains constrained by the limited diversity of high-quality training environments. Generative data augmentation offers a scalable alternative, yet existing methods largely prioritize semantic diversity while imposing limited constraints on local geometry, panoramic wraparound, and cross-node continuity. To address this limitation, we propose a framework that incorporates depth priors, cyclic boundary modeling, and view-adjacency cues into a pretrained diffusion model without additional fine-tuning. Specifically, the uses depth-weighted projection to suppress geometrically implausible content; the jointly denoises cyclic seam regions to preserve left-right continuity; and the reprojects geometry-supported content from neighboring viewpoints to maintain consistency along navigation trajectories. Together, these mechanisms generate visually diverse and spatially coherent panoramic trajectories for VLN training. Experiments across discrete and continuous VLN benchmarks demonstrate consistent improvements in unseen-environment navigation and competitive performance against augmentation methods that rely on large-scale external simulator data. Our code will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.