acceptodds
Under review as a conference paper at ICLR 2027

OS-Themis: A Scalable Critic Framework for Generalist GUI Rewards

Abstract

Reinforcement Learning (RL) has the potential to improve the robustness of GUI agents in stochastic environments, yet its effectiveness critically depends on the quality of the reward signals. Existing methods for constructing such rewards struggle to achieve both scalability and accuracy. We propose OS-Themis, a scalable multi-agent critic framework that decomposes trajectories into verifiable milestones to isolate evidence and audits the evidence chain before producing the verdict. Building on these milestones, we introduce Milestone-Credit Group Relative Policy Optimization (MC-GRPO), leveraging verified milestones for fine-grained intra-trajectory credit assignment in GRPO. We also introduce OmniGUIRewardBench (OGRBench), a holistic multi-platform benchmark for GUI outcome rewards, where OS-Themis consistently outperforms existing reward approaches across models. In online RL, OS-Themis with MC-GRPO improves AndroidWorld performance by up to 9.4%, with MC-GRPO bringing an additional 2.5-2.9% improvement over standard GRPO. We further build a scalable Android online RL infrastructure, where scaling training to 1,024 tasks yields a 10.7% improvement on AndroidWorld, demonstrating the potential of OS-Themis for scalable agent evolution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.