CrossPO: Cross-Trajectory Policy Optimization for User-Centric Agents
Abstract
User-centric LLM agents often need to acquire missing information through interaction before they can complete a task. Learning this behavior is challenging because task outcomes reflect both information acquisition and subsequent execution, leaving trajectory-level rewards too coarse to assess individual information requests. We therefore introduce CrossPO, which learns effective information-seeking behavior by aligning equivalent information requests across trajectories. Within each task's rollout batch, it groups approximately equivalent requests and compares average subsequent returns between trajectories containing and omitting each request class. It shrinks the resulting credit toward zero when either group has limited sample support and uses it to augment GRPO advantages while retaining trajectory-level supervision. This allows equivalent requests to share supervision across dialogue histories, so that a request can receive positive local credit even in an unsuccessful trajectory. Experiments on UserGym, ColBench, and -Bench demonstrate strong overall performance across intent clarification, collaborative coding, and tool use. On ColBench, CrossPO improves pass and success rates over the strongest baseline by up to 5.5 and 6.5 percentage points, respectively. Ablations support the effectiveness of credit assignment and calibration, while behavioral analyses show substantially fewer duplicate requests. Further efficiency analysis shows that CrossPO adds less than 0.5% to per-step training time, making it practical for multi-turn agent training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.