acceptodds
Under review as a conference paper at ICLR 2027

ATLASACT: RETRIEVAL-CONSTRAINED GROUNDING AND STRUCTURE-VALIDATED MEMORY FOR MOBILE GUI AGENTS

Abstract

Multimodal large language models have enabled GUI agents to perform complex mobile tasks, yet agents may identify the intended control correctly while still failing to localize it precisely, and recurring interactions incur repeated model decisions. We propose AtlasAct, a framework that uses accessibility (A11y) structure as a shared executable interface for action grounding and interaction reuse. First, for element-targeted actions, the Actor specifies a target semantically, and retrieval-constrained grounding resolves it against multimodal A11y-node representations. A confidence router directly executes the top-ranked node when retrieval confidence is high and invokes the Grounder only when disambiguation is needed, reducing unnecessary Grounder calls. Second, structure-validated memory maintains Tips as decision guidance and Shortcuts as reusable action sequences. Shortcuts bypass per-step Actor and Grounder calls, while A11y-based entry and execution checks return control to the Actor upon structural mismatch. Experiments on four mobile GUI benchmarks demonstrate improvements in grounding and end-to-end task execution. Memory ablations show that checked Shortcuts reduce average decision steps by 10.4% relative to Tips alone on a fixed task subset, without lowering the observed overall success rate. With Qwen3.8-27B and interaction memory, AtlasAct achieves success rates of 67.67% on AndroidWorld and 95.40% on AndroidDaily, outperforming all evaluated baselines on both benchmarks. Our code is available at https://anonymous.4open.science/r/AtlasAct-A003.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.