SoGUI: Cognition-Conditioned Interaction for Social GUI Agents
Abstract
Progress in vision–language models has made GUI agents increasingly capable, yet most benchmarks provide the interaction goal and primarily evaluate where to act. Social interfaces introduce an earlier challenge: the agent must infer an interaction-relevant semantic state from observed content before a specified policy determines whether and how to interact. We introduce SoGUI, a cognition-conditioned interaction (CCI) formulation that enables separate analysis of semantic inference and policy-conditioned GUI grounding under a controlled policy. We further construct a 3,916-record Theory-Grounded Social-Semantic Cognition Benchmark covering affective valence, visible source-credibility cues, and gain/loss framing, with image-disjoint held-out-platform evaluation. Across three training seeds, SoGUI-SFT-CoG improves Strict Accuracy over the standard cognition-free NoCoG baseline by 6.56/12.37 points on ID/OOD and achieves 94.02%/87.21% E2E-Hit. A prompt-matched NoCoG+Rule control shows that explicit cognition supervision provides additional gains beyond policy-rule prompting alone, improving Strict Accuracy by 2.69/4.73 points on ID/OOD. Error decomposition further shows that once semantic inference is improved, GUI localization becomes the dominant residual bottleneck, particularly under held-out-platform evaluation. These results indicate that semantic cognition and GUI grounding are partially separable capabilities, motivating GUI agents that combine stronger semantic understanding with robust cross-platform grounding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.