AQuaUI: Visual Token Reduction for GUI Agents with Adaptive Quadtrees
Abstract
Large Multimodal Models (LMMs) have recently emerged as promising backbones for Graphical User Interfaces (GUIs) agent models, where high-resolution GUI screenshots are introduced to the prompts at each iteration step with massive amounts of visual tokens. However, these screenshots exhibit highly non-uniform spatial information density: large regions may carry little information and are visually homogeneous, while text and icons require high visual fidelity. Existing approaches to this problem either require additional training or rely on attention-based token compression, ignoring the structured layout and spatial redundancy of GUI screenshots. To fill the gap, this paper proposes __AQuaUI__, a training-free inference-time token reduction method for GUI agent models that utilizes the non-uniform information density in screenshots. AQuaUI constructs an adaptive quadtree on each image and keeps one representative merged token per leaf of the quadtree. It preserves the spatial positions of retained tokens throughout the pipeline to ensure that all position-encoding stages remain consistent. We further propose a conditional quadtree algorithm that utilizes temporal continuity between consecutive screenshots within a single request to refine the current quadtree with previous ones as references to preserve fine-grained regions across static or mildly shifted GUI states. We implement AQuaUI for state-of-the-art GUI agent models on vLLM and conduct experiments on standard grounding and navigational benchmarks. AQuaUI consistently shows improved accuracy-efficiency trade-offs over prior baselines. Notably, on GUI-Owl-1.5-32B-Instruct, AQuaUI reduces average latency by __13.10%__ with __29.52%__ fewer visual tokens while retaining __99.06%__ of full-token performance, suggesting that the spatial redundancy of GUI screenshots can be exploited at inference time without retraining.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.