acceptodds
Under review as a conference paper at ICLR 2027

Diagnosing Personalization Failures in Mobile GUI Agents

Abstract

Unlike conventional GUI agents, personalized GUI agents require user preferences to guide a sequential execution policy over changing GUI states. This creates two challenges for the agentic design: 1) Determining HOW each preference should alter the agent's execution strategy. 2) Determining WHEN each preference should take effect along the interaction trajectory. However, existing benchmarks primarily evaluate overall task success only, making it difficult to evaluate agents’ reasoning ability to determine HOW and WHEN preferences should affect execution. To address this, we introduce PrefDiag-Bench with two key components: 1) a categorization of user preferences into five types, characterizing HOW they affect agent execution; and 2) a stage-wise diagnostic framework with 1,664 probes, tracing WHEN preferences should take effect along the interaction trajectory and where failures occur. Experimentally, we evaluate eleven GUI agent systems, including end-to-end agents and agentic workflows, under two forms of user context: explicit user profiles and behavioral memory logs. We find while existing GUI agents achieve relatively high task completion rates on non-personalized tasks, their performance remains substantially lower on preference-personalized tasks. Even the best-performing model, Seed-2.0-Pro, achieves only 47.17% on process preferences under explicit user profiles. Moreover, agents struggle particularly with Ranking and Process preferences, indicating substantial room for improvement. Further analysis reveals two key capability bottlenecks behind these failures: selecting relevant evidence from user history, and recognizing when user preferences should take effect during execution. Controlled interventions that explicitly provide relevant evidence improve end-to-end task success by up to 7.06%. We hope PrefDiag-Bench can provide a more systematic and comprehensive evaluation of GUI agents’ personalization capabilities, moving beyond simply measuring overall task success.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.