ProtoP: Region-Anchored Visual Token Pruning for Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models translate the reasoning capabilities of large vision-language backbones into continuous robotic control, but their deployment is bottlenecked by the cost of processing dense visual token sequences through Transformer backbones. Token pruning offers a natural remedy, yet we identify a critical failure mode of existing attention-based methods: spatial collapse. At high compression ratios, global Top- election concentrates the retained tokens onto one or two high-attention hotspots - typically the gripper or the currently manipulated object - even though the budget would suffice to cover every region of the scene, leaving the model blind to the rest of it and causing catastrophic task failure. The failure is one of coverage, not of token importance. We introduce ProtoP (Prototype-Preserving Token Pruning), a training-free pruning strategy that structurally prevents spatial collapse. ProtoP partitions the visual attention map into spatially coherent regions around local attention peaks ("prototypes"), then allocates the token budget across regions proportionally to prototype importance, with a one-token-per-region floor that guarantees every semantically distinct region survives. At moderate pruning ratios, where sufficient tokens remain to cover the scene, ProtoP matches the strongest baselines; as the ratio increases and spatial collapse becomes the dominant failure mode, the gap widens sharply, reaching 93.3% (OpenVLA-OFT) and 86.0% () task success at 90% pruning versus 82.2% and 72.8% for the best competing method.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.