False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
Abstract
Self-evolving search agents can construct their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode that we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a corresponding increase in external correctness. A post-hoc reference audit against source evidence shows that co-cheating becomes increasingly severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify each proposal before training. We therefore introduce multi-sample verification (MSV), which queries the same model used in self-evolution three times with the source and three times without it to determine task admission and replace unreliable pseudo-labels. MSV partially reduces false agreement but leaves substantial residual co-cheating and requires six additional labeler generations for every candidate. These limitations motivate CrossFit, our main method. It partitions the proposer's source documents into groups A and B: questions generated from A are scored by an auxiliary solver trained only on B, and vice versa. The resulting cross-fitted agreement determines proposer reward, preventing a same-source pseudo-label from being directly reproduced through the feedback solver while leaving the original solver's update rule unchanged. We evaluate both interventions by rerunning the complete self-evolution loop with Qwen3.5-4B and Qwen3.5-9B. After self-evolution, MSV reduces false-agreement mass from 6.1% to 5.7% on Qwen3.5-4B and from 8.8% to 7.2% on Qwen3.5-9B, whereas CrossFit reduces it to 3.0% and 3.7%, respectively. Replaying identical proposals with source-excluded feedback further reduces false agreement to 0.4% and 0.1%, isolating feedback ancestry from changes in the generated curriculum. Across seven downstream search benchmarks, CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.