acceptodds
Under review as a conference paper at ICLR 2027

Remembering What to Act On: Object-Centric Execution State for Mobile GUI Agents

Abstract

Long-horizon mobile GUI tasks require agents to process multiple data objects across pages and applications. For reliable execution, agents must retain the required values while tracking how task data are organized, what has been done to each object, and which operations the task constraints allow. In this paper, we introduce DOMES, a method for Dynamic Object-centric Maintenance of Execution State, which unifies data retention and progress tracking by binding each object’s structured data to its evolving processing status and task constraints. We train our agent with an online trajectory synthesis pipeline that combines template-based task expansion, automatic environment initialization, and programmatic completion verification. This enables automated collection of successful trajectories that jointly supervise state updates and GUI actions. Considering that state maintenance does not benefit every task, we introduce STR-GRPO, which contrasts rollouts with and without state maintenance on the same task while penalizing non-empty updates. This encourages the agent to maintain state only when it contributes to task success, while improving generalization to unseen tasks. To evaluate agents' ability to complete required data operations and respect task scope, we introduce DataScope, a scalable online mobile benchmark. Its parameterized templates vary the number of target objects and out-of-scope distractors while preserving workflow logic and providing executable environment initialization and final-state verification. Since binary success does not capture partial completion or scope violations, we propose app-level and data-level metrics to measure workflow progress and the completeness and precision of data operations, respectively. Experiments show that our 8B agent achieves a 76.7% success rate on AndroidWorld, outperforming UI-TARS-2-230B by 3.4 percentage points. On DataScope, it consistently outperforms the strongest same-sized baseline across app-level, data-level, and binary task-success metrics. Compared with standard GRPO, STR-GRPO substantially reduces state-update rates while maintaining comparable observed task success. We will release our agent and curated benchmark.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.