acceptodds
Under review as a conference paper at ICLR 2027

ProtoP: Region-Anchored Visual Token Pruning for Vision-Language-Action Models

Abstract

Vision-Language-Action (VLA) models translate the reasoning capabilities of large vision-language backbones into continuous robotic control, but their deployment is bottlenecked by the cost of processing dense visual token sequences through Transformer backbones. Token pruning offers a natural remedy, yet we identify a critical failure mode of existing attention-based methods: spatial collapse. At high compression ratios, global Top- election concentrates the retained tokens onto one or two high-attention hotspots - typically the gripper or the currently manipulated object - even though the budget would suffice to cover every region of the scene, leaving the model blind to the rest of it and causing catastrophic task failure. The failure is one of coverage, not of token importance. We introduce ProtoP (Prototype-Preserving Token Pruning), a training-free pruning strategy that structurally prevents spatial collapse. ProtoP partitions the visual attention map into spatially coherent regions around local attention peaks ("prototypes"), then allocates the token budget across regions proportionally to prototype importance, with a one-token-per-region floor that guarantees every semantically distinct region survives. At moderate pruning ratios, where sufficient tokens remain to cover the scene, ProtoP matches the strongest baselines; as the ratio increases and spatial collapse becomes the dominant failure mode, the gap widens sharply, reaching 93.3% (OpenVLA-OFT) and 86.0% () task success at 90% pruning versus 82.2% and 72.8% for the best competing method.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.