acceptodds
Under review as a conference paper at ICLR 2027

SoGUI: Cognition-Conditioned Interaction for Social GUI Agents

Abstract

Progress in vision–language models has made GUI agents increasingly capable, yet most benchmarks provide the interaction goal and primarily evaluate where to act. Social interfaces introduce an earlier challenge: the agent must infer an interaction-relevant semantic state from observed content before a specified policy determines whether and how to interact. We introduce SoGUI, a cognition-conditioned interaction (CCI) formulation that enables separate analysis of semantic inference and policy-conditioned GUI grounding under a controlled policy. We further construct a 3,916-record Theory-Grounded Social-Semantic Cognition Benchmark covering affective valence, visible source-credibility cues, and gain/loss framing, with image-disjoint held-out-platform evaluation. Across three training seeds, SoGUI-SFT-CoG improves Strict Accuracy over the standard cognition-free NoCoG baseline by 6.56/12.37 points on ID/OOD and achieves 94.02%/87.21% E2E-Hit. A prompt-matched NoCoG+Rule control shows that explicit cognition supervision provides additional gains beyond policy-rule prompting alone, improving Strict Accuracy by 2.69/4.73 points on ID/OOD. Error decomposition further shows that once semantic inference is improved, GUI localization becomes the dominant residual bottleneck, particularly under held-out-platform evaluation. These results indicate that semantic cognition and GUI grounding are partially separable capabilities, motivating GUI agents that combine stronger semantic understanding with robust cross-platform grounding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.