Cobra: Co-Training Bi-Level Observation Routing and Action Policies for GUI Agents
Abstract
The information a GUI agent needs changes from one action to the next, but existing agents typically construct the actor's context with a fixed rule. This can omit useful interface structure or introduce irrelevant information. We model observation selection and action generation as a bi-level partially observable Markov decision process and introduce Cobra (Co-trained Bi-level Routing and Action) to train both policies jointly online. At each step, a lightweight router, GUIRouter, chooses between the screenshot and accessibility (a11y) information and selects the relevant nodes. To learn routing from sparse terminal rewards, we introduce CMA (Counterfactual Modality Advantage), which compares terminal outcomes across routing choices at comparable information states without a learned value network. We analyze when a shared adapter update fails to reduce both modality losses and use two modality-specific LoRA adapters on a frozen shared 8B backbone. With the same initialization, training tasks, and rollout budget, Cobra outperforms the strongest of three fixed-observation online-RL controls by 5.5 percentage points on OSWorld-Verified. It also reduces average per-step a11y tokens by 86% compared with fusing visual input and the a11y tree. Compared with its Qwen3-VL-8B-Thinking base actor, Cobra improves OSWorld-Verified success from 33.9% to 46.3%, with consistent gains on AndroidWorld and Windows Agent Arena.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.