WaLKPruner: Cognitively Inspired Visual Token Pruning for Spatial Understanding
Abstract
Visual token pruning reduces the computational cost of spatial understanding in multimodal large language models. Multi-view inputs contain repeated scene content, but viewpoint changes reveal complementary observations needed to distinguish spatial relationships. Objects overlapping in one view may be separated in another. Feature-guided pruning can preserve the main scene content while discarding spatial details and contextual cues. Geometry-guided pruning maintains diversity across voxels, yet may concentrate tokens in some viewpoints while overcompressing observations from others. Effective pruning must retain representative scene content together with complementary cues along the observation path. Inspired by non-selective global perception and selective processing in human spatial exploration, we introduce WaLKPruner, a training-free framework following the principle of Walk and Look for Key Anchors. Scene abstraction and viewpoint tracking jointly guide token selection, maintaining a compact scene overview while incorporating complementary observations along the ordered observation path. Constrained selective updates limit changes to the initial scene representation while maintaining a fixed token budget. Experiments on three benchmarks show strong performance across token budgets. Efficiency analyses indicate a favorable balance between computational cost and spatial understanding. Quantitative and qualitative analyses reveal balanced token allocation across viewpoints, while ablations support combining scene and path information and retaining observation order. Code and scripts will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.