acceptodds
Under review as a conference paper at ICLR 2027

PACT-GUI: Efficient GUI Agents with Action-Anchored Global–Local History

Abstract

GUI agents for multi-step tasks repeatedly encode high-resolution visual histories, incurring substantial inference costs. While token and KV-cache compression can reduce downstream computation, they leave the cost of visual encoding largely unchanged. However, downsampling historical screenshots can obscure local state changes, while visual redundancy-based pruning may discard seemingly similar regions that remain critical for action prediction. Our analysis shows that GUI histories are globally redundant but locally heterogeneous, and that visually similar content can still affect action prediction. These findings motivate PACT-GUI, a training-free framework that uses executed actions as spatial anchors for history compression. Specifically, before visual encoding, we represent each historical screenshot as a low-resolution global thumbnail paired with a high-resolution crop around the corresponding interaction. To achieve end-to-end acceleration, we pair this representation with template-based structured decoding, which separates the fixed output template from its variable fields, allowing the model to process known syntax in blocks and autoregressively predict only the action type and its parameters. Across four GUI benchmarks, PACT-GUI achieves a $6.62\times$ reduction in mean end-to-end computation, demonstrating substantial efficiency gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.