LLaDA-Nav: Diffusion-Based Action Generation for Visual Navigation
Abstract
Vision-language models (VLMs) provide broad visual-semantic knowledge and promising generalization for embodied AI. To produce navigation actions, autoregressive VLM policies generate action chunks from left to right, conditioning each action on earlier predictions from the same chunk. In our evaluation, however, these policies show limited closed-loop performance in visual navigation. To test whether this feedback pathway contributes to that limitation, we conduct controlled comparisons and prefix interventions that replace model-generated action prefixes with ground-truth ones. Our results suggest that early prediction errors can propagate to later actions, especially under distribution shift. Motivated by this finding, we propose LLaDA-Nav, a visual navigation policy built on a masked diffusion VLM. LLaDA-Nav fully masks the target action block during training and predicts the entire chunk in a single forward pass, eliminating within-chunk target-action reuse. It incorporates previously executed actions as additional context to capture recent robot motion. Across four navigation datasets and unseen HM3D environments, LLaDA-Nav improves open-loop trajectory prediction and achieves 59.03% closed-loop success, exceeding the strongest evaluated baseline by 15.28 percentage points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.