CompGraphRL: Credit Assignment over Compressed Semantic State Graphs for Mobile GUI Agents
Abstract
Online reinforcement learning (RL) in open-ended environments has improved the planning and navigation capabilities of mobile graphical user interface (GUI) agents. However, long-horizon GUI tasks typically yield sparse terminal rewards, creating a severe temporal credit-assignment problem. Existing approaches either assign credit independently within trajectories or detect repeated states through exact observation or environment-state matching. As a result, they fail to identify semantically equivalent states under variations in rendering, dynamic content, and interaction history. We introduce CompGraphRL, a credit-assignment framework based on compressed semantic state graphs. CompGraphRL uses a vision-language model (VLM) to encode each observation jointly with the preceding action and interaction history, merging functionally equivalent states across trajectories into shared graph nodes. We further generalize Generalized Advantage Estimation (GAE) to graph-structured interactions by separately estimating node-level and action-conditioned edge advantages. This design propagates delayed rewards through the graph while providing fine-grained, low-variance signals for policy optimization. Large-scale online RL experiments on MobileGym, with cross-domain evaluation on AndroidWorld and MobileWorld, show that CompGraphRL consistently outperforms sequence-level and graph-based credit-assignment baselines across Qwen3.5 and GUI-OWL-1.5 backbones. With Qwen3.5-27B, it improves MobileGym success from 20.7% to 37.1% and the mean success rate on the two out-of-domain benchmarks from 31.9% to 44.9%. Further analysis shows that a node-merging rate of approximately 15% effectively balances cross-trajectory information sharing and state discriminability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.