SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents
Abstract
Deciding whether a trajectory actually fulfills its instruction determines how we measure computer-use agents on long-horizon graphical-user-interface tasks and how we train them with reinforcement learning (RL). This judgment has long relied on rule-based evaluation. Rule-based evaluation often disagrees with human intention and becomes outdated when an app updates or its online content drifts. Existing model-based judges attempt to address these problems, but their judging accuracy remains limited. We propose the SeekJudge framework, whose four agents, a Condense, a Ground, a Seek and an Analyze agent, reach a verdict through a Seek–Analyze loop over the trajectory. To train a specialized model on densely labeled trajectories, we propose a seed-driven distillation pipeline that expands a few human-labeled seed trajectories into K such trajectories. To evaluate step-level judgments, we build CUAStepBench, a human-annotated benchmark that pairs trajectory verdicts with dense step labels. Beyond accuracy, SeekJudge costs far less than a closed-source large model. We further propose rollout overlap, a reward-server design that reduces the overhead of a reward model in RL training. To our knowledge, SeekJudge is the first model-based reward to match or surpass native rule-based reward in online RL, measured by downstream success rate on held-out RL test goals. In offline judging, SeekJudge-9B also exceeds rule-based evaluation by F1 on AgentRewardBench.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.