AgentJudge:From Real Agent Trajectories to Multi-Granularity Judging and Agent–Judge Co-Evolution
Abstract
Local action quality, overall task performance, and evolving trajectory states are interdependent in agent execution, yet static, outcome-centered judging struggles to capture these signals jointly. We introduce AgentJudge, a framework for multi-granularity judging and agent–judge co-evolution built on real execution trajectories. We collect trajectories from search and coding tasks and derive supervision for local steps, dynamic prefixes, and final outcomes, unifying local action assessment, global outcome evaluation, and execution-state tracking. We adopt progressive SFT–DPO–GRPO training: multi-granularity supervised fine-tuning establishes basic judging capabilities; future-masked retrospective preference learning uses complete trajectories to identify potential misjudgments and constructs revised judgments from evidence visible in the current prefix for preference optimization; reward-based optimization further reinforces structured judging against supervision references. We further design a bidirectional update mechanism in which judge feedback guides agent optimization, while newly generated trajectories supply supervision for subsequent judge updates, forming a loop of trajectory generation, judge learning, and feedback-driven optimization. Experiments show improvements over the base model on general response judging and agent trajectory judging. A small-scale study over two co-evolution rounds further shows joint gains in agent success and judge accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.