Correct Answers, Closed Gates: SolveBind for GUI Agents at the Verification Barrier
Abstract
Real-world websites and applications often deploy CAPTCHAs and other anti-automation mechanisms that use behavioral signals, such as interaction trajectories and timing, rather than outcomes alone, to distinguish human operators from automated agents. We refer to such settings as *protected GUI environments*, where correctly solving a local verification task is insufficient for end-to-end task completion. An agent must also convert solver outputs into interactions accepted by the environment and adapt its execution based on runtime feedback under time constraints. For these environments, we introduce **SolveBind**, a GUI-agent framework that combines *hierarchical solving with runtime supervision*. A hierarchical solving tree decomposes verification into interdependent *Perceive*, *Ground*, *Solve*, *Act*, and *Verify* capabilities. A runtime supervisor grounds action commitments in tool-derived evidence to allow them to be revised as new evidence arrives, and mediates execution under explicit recovery budgets. For verification, across nine CAPTCHA task families and nine backbone models, **SolveBind** achieves the highest success rate on eight backbones, outperforming the strongest baseline by 13.3–34.5 percentage points. For end-to-end task completion, it achieves the highest success rate on all six backbones evaluated on AndroidWorld. In a paired GPT-5.5 analysis, MobileAgent produces correct answers but fails dynamic verification on 13 of 90 tasks; **SolveBind** succeeds on seven of these tasks. With automatic submission retained, removing binding reduces success from 82.22% to 71.56% and increases latency and token usage. Recovery caps reduce execution cost with a small observed success trade-off: removing them raises success from 82.22% to 83.56%, while increasing latency and token usage to 3.7× and 2.4× their full-supervision values, respectively. These results support evidence-grounded execution control while identifying a success–cost trade-off in recovery and failure boundaries of heuristic submission rules.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.