acceptodds
Under review as a conference paper at ICLR 2027

GUI-Drop: State Transition-Aware Visual Token Pruning for Multimodal GUI Agents

Abstract

Vision-language models enable GUI agents to interact with graphical interfaces by encoding high-resolution screenshots into long visual token sequences (e.g., 2,550 tokens for a mobile screenshot). However, most visual tokens are found redundant or irrelevant to action generation, incurring high computational and latency overhead. While existing training-free pruning methods reduce visual tokens, they mostly focus on intra-frame token relevance and redundancy within individual static screenshots. They neglect the logical continuity across GUI state transitions and the complementary cues of UI elements. In particular, they treat localized action feedback identically to static elements, while fragmenting UI elements into disjointed patches that slice complete text and decouple icons from their accompanying labels. To address these limitations, we propose GUI-Drop, a training-free token pruning framework that integrates transition-aware token selection with local feature compensation. We assign higher retention quotas to dynamic regions identified by inter-frame patch similarity and guide token selection by fusing current instruction-to-vision attention with spatially aligned attention signals from prior action generation. To preserve fine-grained local UI evidence, we balance token selection with attention-weighted feature diversity and compensate discarded representations into spatially adjacent retained neighbors. Across diverse online and offline GUI benchmarks and multiple VLM backbones, GUI-Drop generally outperforms existing training-free pruning methods. On AndroidWorld with Qwen3-VL-8B, GUI-Drop outperforms the unpruned baseline at 40% visual token retention. With UI-Voyager, our approach achieves a 59.5% success rate at 20% retention, compared with 43.1% for the strongest baseline, while reducing LLM prefill FLOPs by 41.7% at 40% retention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.