ConGUIBench: Controlled Diagnostics for Multi-Constraint GUI Agents
Abstract
Reliable multi-constraint execution requires graphical user interface (GUI) agents to distinguish goal wording, requested procedure, and realized task state, distinctions that terminal success alone cannot reveal. We introduce ConGUIBench, comprising 118 two- to four-constraint shopping and travel tasks across four web platforms. Controlled Non-Procedural Goal (NPG) and Explicit Procedure (EP) variants, together with state interventions, evaluate goal invariance, procedural controllability, and state adaptivity. Across nine open-weight models on 118 tasks each, five sampled executions per instruction show that reordering equivalent NPG requirements produces 2.8–14.0% of excess cross-order success disagreement beyond same-instruction variability. Under EP, models match the requested order in 79.6–96.0%, yet only 17.6–63.6% of executions both follow the order and succeed. Cross-condition analysis shows that EP sharpens instruction-specific route alignment and increases ISR for eight of nine models, but the completion gains are modest. Agents also respond to visible progress, but suppressing one otherwise valid action lowers paired completion for every model; Oracle state information reduces this loss for most models. An auditor-maintained constraint checklist raises model-equal ISR by 3.7 points under EP and 9.2 points under NPG while also reducing across-run variability on average, highlighting constraint-state tracking as an important reasoning bottleneck. These results expose unstable goal execution, conditional procedural control, and brittle recovery, motivating their separate evaluation in GUI agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.