acceptodds
Under review as a conference paper at ICLR 2027

DON’T LOOK TWICE: LEVERAGING VISUAL EXPERI- ENCE FOR TOKEN-EFFICIENT GUI AGENTS

Abstract

GUI agents powered by multimodal large language models (MLLMs) repeatedly process high-resolution screenshots throughout multi-step interactions, introducing substantial visual computation. Existing efficiency methods primarily reduce redun- dancy within individual screenshots or interaction trajectories, overlooking reusable visual experience accumulated across tasks. We introduce PAVE, a framework that leverages historical interactions for efficient GUI inference. PAVE retrieves relevant experience from previously completed tasks and uses it as a prior for a lightweight controller that adaptively selects visual tokens from the current screen, while keeping the underlying action model unchanged. Experiments on offline GUI benchmarks and online in-the-wild environments show that PAVE reduces visual-token computation by up to 17.4% and 8.3%, while retaining the original task performance. We open source our code and experiment to facilitate further research and application.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.