Semantic Coverage and Temporal Selection for Visual Token Pruning in Vision-Language Navigation
Abstract
Vision-language navigation (VLN) enables embodied agents to follow natural-language instructions in real-world environments. However, dense visual tokens incur substantial computation and inference latency, hindering deployment on resource-constrained robots. Existing VLN token-pruning methods neither explicitly suppress redundancy across historical frames nor exploit instruction semantics during selection. We propose a training-free spatio-temporal visual-token compressor that captures intra-frame semantic diversity and inter-frame novelty for efficient long-horizon inference. For the current frame, low-rank feature grouping identifies complementary semantic groups, while adaptive within-group selection uses feature-space non-maximum suppression to remove redundant tokens without sacrificing semantic coverage. For historical frames, we adopt the same strategy as current frame to construct an expanded candidate pool. Then, we assess each token using current-view relevance, instruction alignment, and visual saliency. Moreover, to reduce temporal redundancy, semi-relaxed temporal alignment uses optimal transport to measure inter-frame novelty against preceding-frame group prototypes. Task-aware token selection uses the combined scores and maximal marginal relevance to iteratively select the final token set. Experiments on standard VLN benchmarks show that our method consistently outperforms competing pruning baselines across compression ratios, achieving a better navigation performance and efficiency trade-off. Real-world quadruped robot experiments further validate its effectiveness for low-latency embodied navigation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.